注册并分享邀请链接,可获得视频播放与邀请奖励。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
加入 January 2024
682 正在关注    150.8K 粉丝
Terminal-Bench-Science discriminates well between models, as well as between different reasoning effort levels within frontier models. Claude Opus 5.5 rises ~38 points from low effort at 24% to xhigh at 62%, alongside a 5x difference in the cost per task. Max effort for Opus 5.5 scores slightly below xhigh at 59%. GPT-6 Sol gains 27 points from low to max effort at ~7.5x the cost per task.
显示更多