Also to be clear - these issues were *flagged by the
@terminalbench team as PRs* (that's how Epoch identified them), and confirmed empirically to affect < 3% of the leaderboard rollouts, i.e. well within reported CIs.
This is not a "broken" benchmark - as is very directly implied by the Epoch report - it's a *continuous* benchmark where
@alexgshaw @ryan_marten and team are earnestly advancing a compelling vision of open benchmarks that are continuously improved over time.