Register and share your invite link to earn from video plays and referrals.

Viv
@Vtrivedy10
applied research @LangChain Labs, prev @awscloud, phd cs @templeuniv
Joined February 2013
1.8K Following    16.4K Followers
good advice for building Evals/Benchmarks (also true for many things in life)… literally just start, build 1 task it’s daunting to have nothing and think about needing to build Terminal Bench but you don’t need that, i promise that having like 5 trusted evals is what you’re really after to start Evals/Environments/Simulations are ridiculously hard, anyone who tells you they’re not is lying to you Literally the frontier research companies in the world spend tons of human hours on Task design, review, and building/refining the system that builds Tasks see Cognition’s Frontier Code ~40 human hours per task Frontier labs spend millions trying to buy these tasks Your first Task will suck, that’s fine, this is research If all your research worked that would be insane but as you keep failing and iterating you’ll build the Eval building muscle for your and your team and that’ll be great because you’ll collectively get better at Evals designed for your most important use cases and that’ll lead to better agents than what copying any existing bench could do for you it’s a marathon, it’s really hard, just get started, publish openly if you can, and reach out to ppl (most want to help anyone trying to do more evals 🚀)
Show more
“the cold start is smaller than people think” - many people complain about designing a comprehensive set of tasks & verifiers. But you don’t NEED this to get started. You need a few good ones & the means to analyze your outcomes/traces so you can update & append; evolve your evals.
Show more