Very exciting to see how much this post resonated with the community. The uncomfortable dialectic is that (1) evals are necessary to build good agents, but also (2) agents are extremely helpful (and necessary) for automating tedious parts of the evals lifecycle. Human attention and labeling simply cannot scale up to the high amount of unstructured info that is agents. Unfortunately nobody has cracked the most perfect and efficient workflow to do evals. We share our thoughts on our current iteration in this post!!