Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
I love waking up to running benchmarks, and checking in on them.
Every morning kinda feels like Christmas.
Here's a quick update on the Fable 5.1 benchmark on my new eval suite, feels good to be able to stump this model a bit.