Register and share your invite link to earn from video plays and referrals.

Search results for DeepSWE
DeepSWE community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including DeepSWE
Grok 4.7 is DeepSWE SoTA at xhigh effort, but take this with a grain of salt.
FrontierCode and DeepSWE scores using API costs as the x-axis instead of the default ouput tokens count
it's not that deepswe just take PTO
eqao is the deepswe of canada
Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.
Show more
0
509
6K
731
Forward to community
Gemini 3.8 Flash on DeepSWE 1.1, scores 73.7%!
0
292
3.6K
133
Forward to community
GLM-5.3-Flash scores 63% on DeepSWE at just $0.24 per task. All evaluations use standard API pricing, not discounted rates.
0
50
1.2K
48
Forward to community
Claude Fable 5 has debuted on DeepSWE Bench with a 66% Pass@1, claiming the #1# spot and edging out GPT-5.5. The result reinforces a broader trend across recent coding benchmarks: strong raw performance combined with consistent reliability and efficiency in real-world software engineering tasks.
Show more
Claude Opus 4.8 has landed on DeepSWE Bench, posting a 58% Pass@1 and taking #2# overall behind GPT-5.5. It continues a broader trend: slightly behind on raw score, but among the most reliable and efficient coding models across recent benchmarks.
Show more
GPT-5.6 Sol Max was already the better deal on DeepSWE v1.1 - The recent price cut widens the gap even further. 🔥 Sol scores 72.7% at $6.47/task, compared with Fable 5 Max at 69.7% and $21.63/task. DeepSWE tests coding agents on 113 original, long-horizon engineering tasks.
Show more