TL;DR Handling a long context doesn't mean a model can actually carry a tedious task through to the end without errors. NVIDIA built Long-Transduction, a new benchmark isolating exactly that sustained-execution ability.
Title: Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
URL:
Points
📉 Accuracy drops 62.8% on average, relatively, when scaling from 4K to 128K tokens
🧮 Four task families — arithmetic, UUID sorting, variable lookup, CSV table transformation — measure sustained execution
📄 Even DeepSeek nailed a perfect document only 41 out of 240 times (17.1%) at 128K tokens
🔀 Just removing stable IDs from the input format drops UUID sorting accuracy by up to 64.3%
🔁 Accuracy consistently declines across every model as local task complexity increases
🤔 Models stay highly self-consistent yet get the wrong answer — they understand the task but can't read the right problem out of the context
🛠️ The authors recommend itemizing inputs with stable IDs, checkpointing output chunks, and splitting work into smaller units
The line that lands hardest: supported context length and model scale don't guarantee reliable, exhaustive execution.
#
AIAgents# #
Benchmark#