Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks.
On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.
Claude Fable 5 has debuted on DeepSWE Bench with a 66% Pass@1, claiming the #1# spot and edging out GPT-5.5.
The result reinforces a broader trend across recent coding benchmarks: strong raw performance combined with consistent reliability and efficiency in real-world software engineering tasks.
Claude Opus 4.8 has landed on DeepSWE Bench, posting a 58% Pass@1 and taking #2# overall behind GPT-5.5.
It continues a broader trend: slightly behind on raw score, but among the most reliable and efficient coding models across recent benchmarks.
GPT-5.6 Sol Max was already the better deal on DeepSWE v1.1 - The recent price cut widens the gap even further. 🔥
Sol scores 72.7% at $6.47/task, compared with Fable 5 Max at 69.7% and $21.63/task.
DeepSWE tests coding agents on 113 original, long-horizon engineering tasks.