So, I finally tested Muse Spark 1.3. 2 real repos, 105 planted bugs, find and fix what you can. Original harness and API.
Big surprise:
Muse Spark 1.3 (max): 33
Fable 5.1 (high): 33
Grok 4.6 (xhigh): 27
Opus 5 (max): 27
Muse Spark 1.3 (high): 19
Meta joined the frontier.
Other effort levels dropping every ~1.5h in this thread 🧵
DeepSeek-V4.1-Flash (high) is in:
max effort: 24/105, $1.80, 42.6 min
high effort: 19/105, $0.31, 26.1 min
For comparision:
Gemini 3.8 Flash (high), 20/105, $9.78, 29.8 min
GLM-5.3 (default): 19/105, $19.73, 66.7 min
Coming next:
deepseek-v4-pro effort=max
deepseek-v4-pro effort=high
Updates are published also on GitHub:
So I tested DeepSeek V4.1 Flash on real tasks. 2 repos, 105 hidden bugs, find and fix what you can.
Opus 5 (max): 27
Grok 4.6 (max): 27
DeepSeek V4.1 Flash (max): 24
GPT-5.6 Luna (xhigh): 23
Opus 5 (high): 21
A really strong model for everyday tasks. And look at the cost:
Opus 5 (max): $51.33
Grok 4.6 (max): $16.96
DeepSeek V4.1 Flash (max): $1.80
GPT-5.6 Luna (xhigh): $2.50
Opus 5 (high): $38.77
It's at the Pareto frontier.
Other effort levels (high) and extra tests for DeepSeek V4 Pro dropping every hour in this thread 🧵
Tried Grok 4.6 on my bug bench an hour after release. 105 hidden bugs in two real repos, judged blind.
Grok 4.5: 17 (+10 non-planted)
Grok 4.6: 27 (+15 non-planted)
Fable 5: 29 (+2 non-planted)
Looks like it may be my new default model. The best combination of time, value, and cost.
4.7 is dropping soon. That may be an even bigger jump.