So I tested DeepSeek V4.1 Flash on real tasks. 2 repos, 105 hidden bugs, find and fix what you can.
Opus 5 (max): 27
Grok 4.6 (max): 27
DeepSeek V4.1 Flash (max): 24
GPT-5.6 Luna (xhigh): 23
Opus 5 (high): 21
A really strong model for everyday tasks. And look at the cost:
Opus 5 (max): $51.33
Grok 4.6 (max): $16.96
DeepSeek V4.1 Flash (max): $1.80
GPT-5.6 Luna (xhigh): $2.50
Opus 5 (high): $38.77
It's at the Pareto frontier.
Other effort levels (high) and extra tests for DeepSeek V4 Pro dropping every hour in this thread 🧵