I think it is unfair to characterize Terminal Bench 4 simply as 'Flawed' or 'Broken' as done in this post: TB4 is a massive community effort. The bugs mentioned are known and public on Github, and the community is working on fixing them.
E.g. saying: 'Terminal Bench 4.0 we found 45.5% of tasks to be broken after reviewing github issues.' is like saying open-source software is BROKEN because some people have found some bugs, raised issues, and the open-source community has not fixed them yet.
Further, in TB4 the current task flaws affect less than 3 percent of rollouts on the leaderboards.
It would be much more useful for a quality audit to find new issues that are not currently known.
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.