Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
More details on the new code quality score that I'm adding into the new version of VulcanBench.
This new frontier is a turning point for benchmarks, excited to be taking the time to make sure we turn in the right direction 🖖