In 13 minutes,
@jeffbarg, Vyshu Khota, and Soroush Khadem walk through how Clay scaled agent evals agents at 300M+ runs a month.
Topics covered:
✅ Their four quadrant eval framework
✅ Why closing the production-to-eval loop is the hardest part
✅ How a data lake and long context changed what agents can do with data