Before their first release, the
@terminalbench team called off a task-writing meetup, figuring there was no way to get more than ten people in a room to talk about the project. Tonight we squeezed ~150 into Laude Lab!
The team covered the state of the bench and
@harborframework, the process behind Terminal-Bench-Science, and how to keep pace as models get better, faster: continuous benchmarks, real-world evals, and long-horizon multi-agent challenges.
@alexgshaw @ryan_marten @StevenDillmann @Mike_A_Merrill @andykonwinski