🚨 DeepSeek V4.1-Flash is kind of insane.
It’s already beating GPT-5.6 Sol on several coding + agent benchmarks:
• DeepSWE: 74.2 vs 73.0
• AutomationBench: 54.8 vs 45.8
• Agents’ Last Exam: 31.8 vs 26.7
• CyberGym: 88.1 vs 84.5
And this is only the Flash model.
V4.1-Pro is still coming.