註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
加入 January 2024
683 正在關注    152.3K 粉絲
Claude Sonnet 5.5 (max) makes large strides on Terminal-Bench, sitting among the top models for both Terminal-Bench 4.0 and Terminal-Bench-Science. In Terminal-Bench 4.0 it scores 64%, a 50 point increase over Claude Sonnet 5 (max), and slightly above 60% for Opus 5.5 and GPT-6 Astra (xhigh). On our leaderboard for Terminal-Bench-Science - a benchmark of agentic terminal use to complete realistic scientific research workflows across domains - it scores 53% and sits behind only GPT-6 Astra and Opus 5.5. Terminal-Bench-Science is not currently included in the Artificial Analysis Intelligence Index.
顯示更多