This matches exactly what we saw building SWE-Together.
For real-world coding tasks, a fixed test suite rarely captures the whole truth. Generated tests can be too strict or low-coverage, too tied to the reference patch, or too shallow about whether the feature actually works. And a failing run can come from environment noise rather than a true model capability gap.
So we added an agentic judge: we freeze a weighted behavioral rubric from the task spec, user intent, and oracle trajectory, then run it in a fresh task sandbox to inspect the patch, workspace, and tests against those goals.
Separating model capability from eval noise takes careful design and audits. The next hard problem is maintenance: benchmarks seem to saturate every few months. We need better ways to refresh tasks reliably and automatically. How should the field approach this?
See more at
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.
We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.