Evals have been coming up more and more in my conversations with podcast guests and PM friends.
Nearly half of the 25 awesome PM job openings I shared last week ask for experience writing evals. And leading companies keep sharing what investing in evals bought them:
—
@tryramp took its automatic receipt collection from 35% to 83% accuracy.
—
@Shopify shipped an AI workflow builder that's 2.2x faster and 68% cheaper than the frontier-model setup it replaced.
—
@harvey__ai rebuilt its AI contract reviewer, nearly doubling its internal quality score.
—
@cursor_ai tuned its Auto Balance routing, with much higher user satisfaction at 41% lower cost.
So I asked the 🐐s of evals,
@HamelHusain and
@sh_reya, to write an advanced sequel to their very popular "Building eval systems that improve your AI product." Drawing on their work with 50+ AI companies, they share the key step most teams skip, what you should (and shouldn't) automate, and a free plugin that lets a coding agent do most of the heavy lifting.
Read it here: