登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
参加 May 2026
281 フォロー中    421 ファン
TL;DR Handling a long context doesn't mean a model can actually carry a tedious task through to the end without errors. NVIDIA built Long-Transduction, a new benchmark isolating exactly that sustained-execution ability. Title: Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability URL: Points 📉 Accuracy drops 62.8% on average, relatively, when scaling from 4K to 128K tokens 🧮 Four task families — arithmetic, UUID sorting, variable lookup, CSV table transformation — measure sustained execution 📄 Even DeepSeek nailed a perfect document only 41 out of 240 times (17.1%) at 128K tokens 🔀 Just removing stable IDs from the input format drops UUID sorting accuracy by up to 64.3% 🔁 Accuracy consistently declines across every model as local task complexity increases 🤔 Models stay highly self-consistent yet get the wrong answer — they understand the task but can't read the right problem out of the context 🛠️ The authors recommend itemizing inputs with stable IDs, checkpointing output chunks, and splitting work into smaller units The line that lands hardest: supported context length and model scale don't guarantee reliable, exhaustive execution. #AIAgents# #Benchmark#
もっと見る