the terminal-bench-science jump was big enough that I went into the repo.
I laughed when I noticed the last commit was authored by Claude. Which is funny, because of course many commits are and obviously there's no meddling here, but also funny in general. They are everywhere. Running the benchmarks, writing the benchmarks, publishing the results.
like, some might say, a new civilization. But let's not use that word just yet.
It’s ahead of both Fable 5 and Opus 5 across our agentic coding and computer use evals.
On Terminal-Bench 4.0 in Claude Code, Fable 5.1 scores 55.8%. Fable 5 scores 42%, and Opus 5 scores 52.3% on the same setup.