Teach your agent taste.
Microsoft and City University of Hong Kong have a very good new paper called The Tasteful Agent, and it explains where a lot of agent budgets go.
You'll recognise the problem if you've left an agent on a long task. It makes a reasonable decision early on, then spends the next few hours building on it, and the whole run goes nowhere.
On a long task, an agent keeps choosing what to try next: which bug fix to attempt, which version of the code to build on and whether to stick with an approach or start again. All options usually look sensible so the agent finds out it picked wrong after it has spent most of your money on the wrong path.
The authors call the ability to choose well at these moments "taste," and they built a way to test it.
They went through 2,677 coding runs and 1,132 research runs, looking for points where two attempts took different paths and one turned out better.
They turned each of those moments into a question: given what the agent knew at the time, which way should it go?
Of 4,657 candidate forks, 502 made it through the filters. Human reviewers agreed with the labels 98.8% of the time.
They tested 14 frontier models, including GPT-5.6 Sol, Claude Opus 5, Grok and DeepSeek. To get a point, a model had to choose the same answer with the options shown in either order. Guessing scores 25%.
Four things stood out.
The best model, GPT-5.6 Sol, scored 59.7%. It misses about four of these calls in ten, and on a long task each miss sends hours of work down the wrong path.
When the deciding clue was in what the agent had already seen, models averaged 62.3%. When the clue only appeared after more work, they fell to 21%, below guessing. Long tasks are made of exactly that second kind of decision.
More thinking time didn't help. The researchers tested two models at three reasoning levels. Neither improved, and both burned the most tokens on the decisions they got wrong most often. The bigger reasoning budget went to the forks where it bought nothing.
Taste can be trained though. They took a 27B open model, Qwen3.6, and trained it on how past forks had turned out. Its accuracy on unseen tasks rose by 17.9 percentage points.
Then they tested it on real coding work. A coding agent got advice from that small model before tackling 41 SWE-bench Pro tasks it hadn't seen. Its success rate went from 14.6% to 33.7%, more than double. Perfect advice would have taken it to 39%.
A 27B open model that studied old runs more than doubled what the agent could finish.
The pitch from the big labs has been to pay for a bigger model and let it think longer. This paper found that on the decisions that sink long tasks, longer thinking changed nothing and the top models still miss four in ten.
Companies are starting to hand agents work that runs for hours with nobody watching. Every one of those runs passes through forks like these, and the agent picks without knowing what comes next.
If you run agents, keep the failed runs.
They're a record of every fork that went wrong, and this paper used runs like them to teach an agent to pick better next time.
もっと見る