Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
Stop everything, benchmark Grok 4.6.
And yeah, quite a few steps to get these benchmarks setup. I typically use Fable 5 Low effort as my orchestrator for these.
As you can see, lots of updates need to be made to make sure pricing and effort are benchmarked correctly when a new model comes out.