Terminal-Bench-Science discriminates well between models, as well as between different reasoning effort levels within frontier models. Claude Opus 5.5 rises ~38 points from low effort at 24% to xhigh at 62%, alongside a 5x difference in the cost per task. Max effort for Opus 5.5 scores slightly below xhigh at 59%. GPT-6 Sol gains 27 points from low to max effort at ~7.5x the cost per task.
顯示更多