TL;DR Handling a long context doesn't mean a model can actually carry a tedious task through to the end without errors. NVIDIA built Long-Transduction, a new benchmark isolating exactly that sustained-execution ability.
Title: Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
URL:
Points
๐ Accuracy drops 62.8% on average, relatively, when scaling from 4K to 128K tokens
๐งฎ Four task families โ arithmetic, UUID sorting, variable lookup, CSV table transformation โ measure sustained execution
๐ Even DeepSeek nailed a perfect document only 41 out of 240 times (17.1%) at 128K tokens
๐ Just removing stable IDs from the input format drops UUID sorting accuracy by up to 64.3%
๐ Accuracy consistently declines across every model as local task complexity increases
๐ค Models stay highly self-consistent yet get the wrong answer โ they understand the task but can't read the right problem out of the context
๐ ๏ธ The authors recommend itemizing inputs with stable IDs, checkpointing output chunks, and splitting work into smaller units
The line that lands hardest: supported context length and model scale don't guarantee reliable, exhaustive execution.
#
AIAgents# #
Benchmark#