I wanted to share a bit more about the benchmark results of our first voice model comparison with VulcanBench.
We asked Grok Voice Think Fast 2.0 and GPT Realtime the same 200 questions, typed vs spoken.
Here's the results:
Grok: 99.0% text → 95.7% audio (+3.3 pp)
GPT Realtime: 97.5% → 93.5% (+4.0 pp)
Real, but small. And here's a little thread for anyone that wants to dive deeper into the details.
A voice model benchmarking thread 🧵
Okay, the benchmark I've been waiting all week to run is now running.
Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5
Low and Medium effort levels.
All on the new v3 eval suite. Can't wait to wake up in the morning to see the results.
Live long and benchmark 🖖
Thanks to Fable, VulcanBench v2 is now live on Github.
The goal with this release was a big one, so I wanted to wait for Fable to come back to do it.
I kept paying $100+ to run my benchmarks where every frontier model scored 98-100%.
So I rebuilt it with Fable.
VulcanBench v2 is 10 tasks pulled from real merged PRs in flask, aiohttp, sqlglot, click, and chi (Python + Go).
Everything merged after model training cutoffs so nothing is memorized, all graded by deterministic hidden tests.
Oh and one really nice addition that seemed like the best way to optimize my costs...prompt caching in the harness, a full two-model run now costs $7.36 instead of $100+.
Keep calm and benchmark on 🖖