Register and share your invite link to earn from video plays and referrals.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
Joined May 2026
280 Following    416 Followers
🎯 A new way to measure an agent's judgment quality on long-horizon tasks, with no expert annotation required. Title: The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks URL: 📖 Overview This work builds Taste-Bench, a 502-question benchmark that evaluates the decisions an agent makes mid-task (which hypothesis to test, which implementation to build on). It also shows this judgment can be trained through distillation. 🔥 The problem it solves Existing benchmarks only measure whether a task was completed, not the quality of the choices made along the way. A bad decision looks perfectly reasonable when it's made, and its cost only surfaces after the agent has burned most of its budget. Human grading needs deep domain expertise and doesn't scale. 🧪 Methodology The key insight is that the later part of a trajectory is hindsight evidence for its earlier decisions. ・Points where parallel attempts at the same task diverge and only one succeeds (mistakes the agent never notices) ・Points inside a single run where the agent hits failure and recovers (mistakes the agent self-corrects) Both kinds of forks are mined automatically, everything after the fork is hidden, and the model picks between two directions. Each question is scored in both candidate orders and counts as correct only if both are right, so random guessing scores 25%. 📊 Results ・Across 14 frontier models, the best is GPT-5.6 Sol at 59.7% — nowhere near ceiling for a binary choice ・Accuracy falls sharply as the deciding evidence moves further out, bottoming at 21.0%, below random ・Maxing out the reasoning budget moves accuracy by only −0.2 to +2.2 points ・Correlation with SWE-bench Verified is just r=0.63; models within 4 points there differ by 10.7 here ・Injecting a distilled student's advice lifts real task success from 14.6% to 33.7% #AIAgents# #Benchmarks#
Show more