Register and share your invite link to earn from video plays and referrals.

terminalbench
@terminalbench
4 Following    749 Followers
Before their first release, the @terminalbench team called off a task-writing meetup, figuring there was no way to get more than ten people in a room to talk about the project. Tonight we squeezed ~150 into Laude Lab! The team covered the state of the bench and @harborframework, the process behind Terminal-Bench-Science, and how to keep pace as models get better, faster: continuous benchmarks, real-world evals, and long-horizon multi-agent challenges. @alexgshaw @ryan_marten @StevenDillmann @Mike_A_Merrill @andykonwinski
Show more
Terminal-Bench meetup tonight *with* swag. Register if you haven’t already! Benchmarks, RL environments, new features, and state of the union
fast track to get model labs to care about the capabilities you care about: contribute a task to Terminal-Bench if you have built a benchmark around a specific use case, DM me and we can collaborate on a TB task for the next release
Show more
We're hosting a meetup! Come meet the community behind the benchmark, celebrate recent releases, and join our discussions about the future of agent evaluation.
Terminal-Bench-Science 0.1 is the #1# featured benchmark on @AnthropicAI’s new Claude Fable release 🚀
Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
Show more
0
128
5.5K
449
Forward to community
Terminal-Bench 4.0 out now!
We've pushed a version update to the Terminal-Bench dataset and leaderboard. Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
Show more
Announcing Terminal-Bench-Science!
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%. 1/n 👇
Show more
Congrats to for the strong performance on Terminal-Bench 3.0! One of the biggest pieces of feedback we have gotten for TB3 is to increase the timeouts. We calibrated timeouts against frontier models during development, but inference speed can still be a confounder on some of the tasks. Terminal-Bench numbers on the GLM-5.3 model card are reported with increased timeouts (likely for this reason). Look out for Terminal-Bench 4.0 releasing soon with increased timeouts, other task improvements, and a handful of new tasks.
Show more
Grok 4.6 now on the Terminal-Bench 3.0 Leaderboard!
Grok 4.5 is SOTA on TB2.1... at reward hacking In all seriousness, even after zeroing out reward hacks, it is #4# on the TB2.1 leaderboard and lands on the Pareto for both cost and speed. (charts and reward hacking links in 🧵)
Show more
GPT‑5.6 Sol sets a new state of the art on Terminal‑Bench 2.1, which tests complex command-line workflows requiring planning, iteration, and tool coordination.
0
108
3.9K
233
Forward to community
Can agents build complete projects that deliver real value? We’re launching Terminal Bench Challenges: 3 unsolved tasks which could make a real impact on the open source community if solved. These tasks provide a testing ground for optimizations both on the model and harness level on our continuous leaderboard for each task.
Show more