good advice for building Evals/Benchmarks (also true for many things in life)…
literally just start, build 1 task
it’s daunting to have nothing and think about needing to build Terminal Bench
but you don’t need that, i promise that having like 5 trusted evals is what you’re really after to start
Evals/Environments/Simulations are ridiculously hard, anyone who tells you they’re not is lying to you
Literally the frontier research companies in the world spend tons of human hours on Task design, review, and building/refining the system that builds Tasks
see Cognition’s Frontier Code ~40 human hours per task
Frontier labs spend millions trying to buy these tasks
Your first Task will suck, that’s fine, this is research
If all your research worked that would be insane but as you keep failing and iterating you’ll build the Eval building muscle for your and your team
and that’ll be great because you’ll
collectively get better at Evals designed for your most important use cases
and that’ll lead to better agents than what copying any existing bench could do for you
it’s a marathon, it’s really hard, just get started, publish openly if you can, and reach out to ppl (most want to help anyone trying to do more evals 🚀)
“the cold start is smaller than people think” - many people complain about designing a comprehensive set of tasks & verifiers. But you don’t NEED this to get started. You need a few good ones & the means to analyze your outcomes/traces so you can update & append; evolve your evals.