登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Viv
@Vtrivedy10
applied research @LangChain Labs, prev @awscloud, phd cs @templeuniv
参加 February 2013
1.8K フォロー中    16.4K ファン
good advice for building Evals/Benchmarks (also true for many things in life)… literally just start, build 1 task it’s daunting to have nothing and think about needing to build Terminal Bench but you don’t need that, i promise that having like 5 trusted evals is what you’re really after to start Evals/Environments/Simulations are ridiculously hard, anyone who tells you they’re not is lying to you Literally the frontier research companies in the world spend tons of human hours on Task design, review, and building/refining the system that builds Tasks see Cognition’s Frontier Code ~40 human hours per task Frontier labs spend millions trying to buy these tasks Your first Task will suck, that’s fine, this is research If all your research worked that would be insane but as you keep failing and iterating you’ll build the Eval building muscle for your and your team and that’ll be great because you’ll collectively get better at Evals designed for your most important use cases and that’ll lead to better agents than what copying any existing bench could do for you it’s a marathon, it’s really hard, just get started, publish openly if you can, and reach out to ppl (most want to help anyone trying to do more evals 🚀)
もっと見る
“the cold start is smaller than people think” - many people complain about designing a comprehensive set of tasks & verifiers. But you don’t NEED this to get started. You need a few good ones & the means to analyze your outcomes/traces so you can update & append; evolve your evals.
もっと見る