We audited 15 benchmarks and labeled 9 flawed:
- In Terminal Bench 4.0 we found 45.5% of tasks to be broken after reviewing github issues.
- In HLE, 46% of the 48 questions we randomly sampled were broken.
- In DeepSWE 1.1 we found a bug that can break grading for every task.