註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
加入 January 2024
682 正在關注    150.8K 粉絲
As with our specification for Terminal-Bench 4.0, we run Terminal-Bench-Science with the mini-swe-agent harness across all models. This keeps comparisons like-for-like, and in our testing often retains comparable evaluation performance to first-party harnesses like Codex and Claude Code. We're excited to keep updating the leaderboard as the frontier shifts. The Terminal-Bench-Science team’s call for contributions for 0.2 is open until October 5 at Explore the full results at
顯示更多