注册并分享邀请链接,可获得视频播放与邀请奖励。

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
加入 May 2026
281 正在关注    421 粉丝
TL;DR Handling a long context doesn't mean a model can actually carry a tedious task through to the end without errors. NVIDIA built Long-Transduction, a new benchmark isolating exactly that sustained-execution ability. Title: Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability URL: Points 📉 Accuracy drops 62.8% on average, relatively, when scaling from 4K to 128K tokens 🧮 Four task families — arithmetic, UUID sorting, variable lookup, CSV table transformation — measure sustained execution 📄 Even DeepSeek nailed a perfect document only 41 out of 240 times (17.1%) at 128K tokens 🔀 Just removing stable IDs from the input format drops UUID sorting accuracy by up to 64.3% 🔁 Accuracy consistently declines across every model as local task complexity increases 🤔 Models stay highly self-consistent yet get the wrong answer — they understand the task but can't read the right problem out of the context 🛠️ The authors recommend itemizing inputs with stable IDs, checkpointing output chunks, and splitting work into smaller units The line that lands hardest: supported context length and model scale don't guarantee reliable, exhaustive execution. #AIAgents# #Benchmark#
显示更多