Great paper from Amazon.
In discusses when not to trust LLM judges for agent evaluation.
(bookmark it)
A common way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. This paper from Amazon shows that gate fails in two specific ways.
1. Satisfaction does not track success. 57.5% of conversations the raters marked satisfied had failed the customer's task.
2. Close calls go wrong. The ranking holds across agents of very different ability, but among near-equal agents the gate picks the lower-reward one on 31% of pairs, compared with under 1% for pairs far apart.
The study covers 25 agents from six providers on tau2-bench and SimulatorArena. Judges also favored agents from their own model family.
The fix is cheap. A judge-free completion bit catches truncation regressions, and the judge is trusted only after calibration against a verifiable reward.
Paper: