登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
参加 January 2024
682 フォロー中    152K ファン
Claude Sonnet 5.5 (max) makes large strides on Terminal-Bench, sitting among the top models for both Terminal-Bench 4.0 and Terminal-Bench-Science. In Terminal-Bench 4.0 it scores 64%, a 50 point increase over Claude Sonnet 5 (max), and slightly above 60% for Opus 5.5 and GPT-6 Astra (xhigh). On our leaderboard for Terminal-Bench-Science - a benchmark of agentic terminal use to complete realistic scientific research workflows across domains - it scores 53% and sits behind only GPT-6 Astra and Opus 5.5. Terminal-Bench-Science is not currently included in the Artificial Analysis Intelligence Index.
もっと見る