Frontier-Bench picks up where Terminal-Bench left off. Why the new name?
First, it evaluates capabilities at the frontier beyond agentic coding: finance, music, biology, hardware design, etc.
Second, it evolves in lockstep with the frontier.
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work.
Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort.
Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%