๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
281 ํŒ”๋กœ์ž‰ ์ค‘    421 ํŒฌ
TL;DR Handling a long context doesn't mean a model can actually carry a tedious task through to the end without errors. NVIDIA built Long-Transduction, a new benchmark isolating exactly that sustained-execution ability. Title: Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability URL: Points ๐Ÿ“‰ Accuracy drops 62.8% on average, relatively, when scaling from 4K to 128K tokens ๐Ÿงฎ Four task families โ€” arithmetic, UUID sorting, variable lookup, CSV table transformation โ€” measure sustained execution ๐Ÿ“„ Even DeepSeek nailed a perfect document only 41 out of 240 times (17.1%) at 128K tokens ๐Ÿ”€ Just removing stable IDs from the input format drops UUID sorting accuracy by up to 64.3% ๐Ÿ” Accuracy consistently declines across every model as local task complexity increases ๐Ÿค” Models stay highly self-consistent yet get the wrong answer โ€” they understand the task but can't read the right problem out of the context ๐Ÿ› ๏ธ The authors recommend itemizing inputs with stable IDs, checkpointing output chunks, and splitting work into smaller units The line that lands hardest: supported context length and model scale don't guarantee reliable, exhaustive execution. #AIAgents# #Benchmark#
๋” ๋ณด๊ธฐ