Okay, my Muse Spark 1.2 benchmark with
@VulcanBench is done, and holy moly, this was a wild one.
I benchmarked it on Eval Suite 3, with the same rules I give every model, and it just couldn't get a good chunk of tasks solved in the time limit.
As a quick reminder, VulcanBench is focused on all real engineering tasks, things swe's would give to a model during regular daily work, no weird math puzzles or exotic architecture.
And every model is given the same time constraints, scaled by task difficulty. This is how engineering leaders like me make decisions, we can't have models that solve something in an infinite amount of time, when another model can solve it 10x faster.
Grok 4.5 High remains at the top of the leaderboard, you can see the full leaderboard here:
For more details on the Muse Spark 1.2 benchmark, you can see the detailed report here: