Good question. The main secret sauce is... you :) Basically we believe that evals are all about pairing domain experts with systems that allow them to encode their taste and judgement and put this all together in a workflow where you're improving your agent by seeing what's going wrong, using deterministic (where possible) evaluators to capture those failure points, and then using the replay etc to make sure that you've actually fixed things at the root. We're not really at the point where you can just automate evals fully without humans being involved (see
@HamelHusain's recent post on some of the ways that can go wrong), but for sure tools (like coding agents, or like Kitaru) can help make this process as painless as possible.