SWE-bench Verified is contaminated. OpenAI just published the proof.
Top models: 70%+ on Verified, ~23% on SWE-bench Pro.
All frontier models can reproduce original fixes from memory. 59% of hard tasks have flawed tests.
The benchmark everyone was citing? Meaningless now.