DeepSeek V4.1 Flash dropped and the benchmarks say it beats Opus 5 and GPT 5.6 Sol on DeepSWE.
I tested it yesterday. It does not.
This is the second Flash model in eight days to "beat" Opus 5 on DeepSWE. Gemini 3.8 Flash last week. DeepSeek today. Neither one comes close in a real codebase.
The whole point of a benchmark is to tell you how a model performs before you spend your money and your time on it. That is the entire job.
Instead the labs train to the test, post the chart, and let you find out the hard way.
Benchmarks are broken.