We’re investing heavily in evals at
@sentry. It’s early but we’re seeing a culture change from shipping vibes to shipping with clear data using our evals platform.
Will share more in a future blog post, but some stats from our devs in the past month:
- ~1800 eval runs triggered directly on PRs
- 1500 eval scenarios ran across 17 AI surfaces in our product
- Cost/token reductions of ~20-50% across a few of our surfaces
- Eval passing improvements of ~10-20% from some core “code mode” changes launching soon (shoutout
@grichadev)
All this from essentially 0 back when we started this effort in May.