註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Fred Oliveira 🧠
@f
working towards better AI futures. AI safety at @snyksec. Organizing @lisbonai_. Data and capital at @capitalfactory. Prev: @gumroad @techcrunch @oreillymedia
加入 September 2006
922 正在關注    133.9K 粉絲
the terminal-bench-science jump was big enough that I went into the repo. I laughed when I noticed the last commit was authored by Claude. Which is funny, because of course many commits are and obviously there's no meddling here, but also funny in general. They are everywhere. Running the benchmarks, writing the benchmarks, publishing the results. like, some might say, a new civilization. But let's not use that word just yet.
顯示更多
It’s ahead of both Fable 5 and Opus 5 across our agentic coding and computer use evals. On Terminal-Bench 4.0 in Claude Code, Fable 5.1 scores 55.8%. Fable 5 scores 42%, and Opus 5 scores 52.3% on the same setup.
顯示更多