What if you don't need Astra, or Sol, or Terra, or Luna, for a lot of your daily coding tasks, and GPT 5.5 does them just fine? Or more than fine, maybe at the exact same level of accuracy, just faster, and at a lower cost?
And what if you don't need GPT 5.5 High or xHigh, but actually medium, 80% of the time?
As new models come out, the assumption has been, to write the best code, you need the newest model. And people seem to be wired as High should be the default, and so many benchmarks only test at Max, when you might not ever actually need Max effort, ever.
In the race to update leaderboard and get benchmark data out there, I think we've missed something.
Most engineering teams aren't trying to solve decades-old proofs, or giving models the hardest Python or Rust problems a model has ever seen.
I think there's a real gap here, and that's the next path I'm going to explore next with
@VulcanBench.
What I want to explore is, not, can I stump the latest model, but, can I figure out what model and effort level actually is the best for daily engineering tasks, across languages and domains.
I just finished a preliminary test with GPT 5.5 Medium, and I'm pretty blown away with what I'm seeing.
Running another test now with Opus 4.6, I think there's something interesting here.
More to come 🖖