Register and share your invite link to earn from video plays and referrals.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
Joined March 2020
63 Following    2.3K Followers
We are going back to the future with our next benchmark report w/VulcanBench. The next test will be comparing GPT 5.5 across effort levels to Luna, Terra, and Sol, across effort levels. More to come 🖖
Show more
What if you don't need Astra, or Sol, or Terra, or Luna, for a lot of your daily coding tasks, and GPT 5.5 does them just fine? Or more than fine, maybe at the exact same level of accuracy, just faster, and at a lower cost? And what if you don't need GPT 5.5 High or xHigh, but actually medium, 80% of the time? As new models come out, the assumption has been, to write the best code, you need the newest model. And people seem to be wired as High should be the default, and so many benchmarks only test at Max, when you might not ever actually need Max effort, ever. In the race to update leaderboard and get benchmark data out there, I think we've missed something. Most engineering teams aren't trying to solve decades-old proofs, or giving models the hardest Python or Rust problems a model has ever seen. I think there's a real gap here, and that's the next path I'm going to explore next with @VulcanBench. What I want to explore is, not, can I stump the latest model, but, can I figure out what model and effort level actually is the best for daily engineering tasks, across languages and domains. I just finished a preliminary test with GPT 5.5 Medium, and I'm pretty blown away with what I'm seeing. Running another test now with Opus 4.6, I think there's something interesting here. More to come 🖖
Show more