注册并分享邀请链接,可获得视频播放与邀请奖励。

Fred Oliveira 🧠
@f
working towards better AI futures. AI safety at @snyksec. Organizing @lisbonai_. Data and capital at @capitalfactory. Prev: @gumroad @techcrunch @oreillymedia
加入 September 2006
922 正在关注    133.9K 粉丝
the terminal-bench-science jump was big enough that I went into the repo. I laughed when I noticed the last commit was authored by Claude. Which is funny, because of course many commits are and obviously there's no meddling here, but also funny in general. They are everywhere. Running the benchmarks, writing the benchmarks, publishing the results. like, some might say, a new civilization. But let's not use that word just yet.
显示更多
It’s ahead of both Fable 5 and Opus 5 across our agentic coding and computer use evals. On Terminal-Bench 4.0 in Claude Code, Fable 5.1 scores 55.8%. Fable 5 scores 42%, and Opus 5 scores 52.3% on the same setup.
显示更多