Loading up the APIs for my next
@VulcanBench run.
Put a lot into this one, v3 eval suite with 23 new evals. Broader language coverage. And some really deep dives into making sure this doesn't have any of the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes.
Still, I'm new to benchmarking, so there's always a chance that I waste like $150 here. Hoping that's not the case, feels like I've tried to think of every edge case.
This is the comparison I've been so excited to do:
Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5
All on Low and Medium effort, across real tasks that represent the work real engineering teams will be doing daily. No puzzles, no weird math problems, and looking not just at accuracy, but token use, cost, and time.
In aggregate this will be a total of 138 runs, will wake up on Saturday to see the results.