๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
270 ํŒ”๋กœ์ž‰ ์ค‘    316 ํŒฌ
Do all your agent's LLM calls really need a frontier model? NVIDIA's Switchyard ran the numbers โ€” and the results are surprising. Title: Switchyard Agent Routing Benchmark URL: TL;DR Evaluated 145 multi-step agentic tasks (averaging 6.3 LLM calls each). 93% of calls were handled by a smaller model, achieving 74% cost reduction with only a 6-point accuracy drop. Key Points ๐ŸŽฏ Frontier model needed for just 7% of calls Out of all LLM calls, Claude Opus 4.8 was required for only 7%. The remaining 93% were handled by Nemotron 3.5 Lightning. ๐Ÿ’ฐ 74% cost reduction Per-task cost dropped from $0.092 (Opus alone) โ†’ $0.026 (routed) โ†’ $0.006 (Lightning alone), while maintaining 80% accuracy. ๐Ÿ“Š The surprising cost breakdown Despite handling only 7% of calls, the frontier model consumed 68.4% of total spend. The judge model itself added another 21.2% of routed costs. ๐Ÿ”ข The formula for routing ROI "Minimum offload rate = judge cost / (expensive model cost - cheap model cost)" โ€” if the price gap is small, routing may not pay off. โšก Two deployment options Run NVIDIA Switchyard as a standalone proxy server, or embed it as middleware inside LangChain's Deep Agents framework. Note: the eval suite had relatively easy tasks, so the benefit of routing could be even larger with harder workloads. Still, highly practical for agents that issue many calls per task. #AIAgents# #LLMCost#
๋” ๋ณด๊ธฐ