Lots of good examples in these reports.
For example: in DeepSWE v1.1, the verifier discards the agent's changes to some test files. But the model isn't told this will happen! Which leads to all sorts of 'irrelevant' failures...
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.