When an agent takes more actions but makes no real progress, this is a useful early-warning sign that it’s heading towards failure.
Across 13K+ OfficeQA runs, agents that ultimately answered incorrectly took up to 50% more steps, consumed ~40% more compute, and incurred ~40% higher cost per episode than agents that reached the correct answer.
TL;DR: Failed trajectories don't just take longer. They waste significantly more compute, tokens, and money.