Grok 4.5 just ranked #
1# on VulcanBench, outperforming Claude Fable 5, GPT-5.6 Sol and Kimi K3 across 23 real-world merged PRs in Python, Rust, TypeScript, JavaScript and Go
Insane efficiency:
• 21/23 tasks solved - 91.3%
• Best result across all 11 configurations tested
• Just $0.32 per solved task
Claude Fable 5 and GPT-5.6 Sol peaked at 20/23. Kimi K3 needed a two-hour extended budget just to reach the same score and cost $1.37 per solved task
Grok 4.5 did not just win on accuracy. It won on economics too
Grok’s coding efficiency is insane