๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
279 ํŒ”๋กœ์ž‰ ์ค‘    414 ํŒฌ
What if letting an agent evolve its own workflow just meant it got better at gaming the test? This paper tackles exactly that overfitting problem. Title: RRSI: Regularized Recursive Self-Improvement of Agent Harnesses URL: โ“ What's a harness? It's everything wrapped around the frozen LLM: prompts, control flow, tooling, memory, context management. The same model can perform very differently depending on this design. โ“ Why does self-improving it overfit? Three failure modes show up: fitting too tightly to the evolve-set benchmark, chasing noise, and letting complexity pile up unchecked. ๐Ÿ’ก How does RRSI fix it? It applies classic ML regularization (L0, L1, L2) to harness evolution: an annealed edit budget limits changes per round, benchmark-specific proposals get screened out at selection, and unproductive components get pruned. ๐Ÿ’ก What's the payoff? Across 8 benchmarks, RRSI beats baselines by up to 22.9% on unseen tasks, while using 30% fewer tokens. It feels like a genuinely grounded step toward agents that can safely improve their own workflow in production. #AIAgents# #SelfImprovement#
๋” ๋ณด๊ธฐ