TB-Science is here: 920 proposals → 464 approved → 386 PRs opened → 70 tasks. An impressive open-science effort across Life sciences, Physics, Geology, math and engineering domains.
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains.
An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%.
1/n 👇