Register and share your invite link to earn from video plays and referrals.

Datacurve
@datacurve
Research and data to advance frontier models.
17 Following    11K Followers
Gemini 3.8 Flash achieves 73.7% on DeepSWE. It pushes the frontier with an impressive 8.2% increase over Gemini 3.7 Flash, at the same cost but with more steps and output tokens.
Show more
Gemini 3.8 Flash on DeepSWE 1.1, scores 73.7%!
0
292
3.6K
133
Forward to community
Gemini 3.7 Flash debuts at 65.5% on DeepSWE. It delivers substantial improvements over 3.6 Flash, scoring +18.8% higher while costing less than half as much per task.
Datacurve is at @aiDotEngineer! Come by our booth in the expo hall to learn more about DeepSWE (and snag some merch). Plus, join us on Thursday @ 10:30am for our talk.
We've updated the Artificial Analysis Coding Agent Index, replacing SWE-Bench Pro with Datacurve's DeepSWE benchmark - the swap lifts Codex with GPT-5.5 (xhigh) above Claude Code with Opus 4.8 (max), while the newly released Claude Fable 5 (max) in Claude Code debuts at the top DeepSWE, built by @datacurve, writes its tasks from scratch rather than adapting them from public GitHub issues or pull requests, so no model has seen the solutions during training. That matters because SWE-Bench Pro, the benchmark it replaces in our Coding Agent Index, had grown gameable, with some models recovering the fix from the repository's commit history instead of solving the task. The swap reorders the index: Codex with GPT-5.5 (xhigh) rises from 65 to 76, overtaking Claude Code with Opus 4.8 (max) at 73. Claude Code with Fable 5 (max), which enters directly on the refreshed index, leads at 77. SWE-Bench Pro had been flattering some combinations and penalizing others. More below.
Show more
0
112
2K
181
Forward to community
Opus 4.8 is now on DeepSWE. On the default high thinking effort, it scores 6% higher than Opus 4.7 xhigh, while also lowering average cost per task.
0
87
1.8K
116
Forward to community