Not Diamond just made coding agent evals way more realistic.
They simulate actual user sessions with long context, cache expiry, and messy interactions.
Their router gets Opus xhigh-level quality at 20–60% lower cost.
this is a pretty big deal.
Today we’re releasing our methodology for evaluating model routing with interactive benchmarks, which represent agent cost accumulation better than static benchmarks do.
Across leading benchmarks, we achieve Pareto-dominance, exceeding Opus xhigh quality at 20–80% lower cost.