Super cool work! This also explains why some benches have an "artificial ceiling".
For example, on DeepSWE, Epoch found 23 false negatives out of 131 tasks. This means a ceiling of roughly 79.6%, which actually matches the current (almost) saturated high score of ~74%
顯示更多
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
顯示更多