註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
加入 May 2026
280 正在關注    413 粉絲
What if letting an agent evolve its own workflow just meant it got better at gaming the test? This paper tackles exactly that overfitting problem. Title: RRSI: Regularized Recursive Self-Improvement of Agent Harnesses URL: ❓ What's a harness? It's everything wrapped around the frozen LLM: prompts, control flow, tooling, memory, context management. The same model can perform very differently depending on this design. ❓ Why does self-improving it overfit? Three failure modes show up: fitting too tightly to the evolve-set benchmark, chasing noise, and letting complexity pile up unchecked. 💡 How does RRSI fix it? It applies classic ML regularization (L0, L1, L2) to harness evolution: an annealed edit budget limits changes per round, benchmark-specific proposals get screened out at selection, and unproductive components get pruned. 💡 What's the payoff? Across 8 benchmarks, RRSI beats baselines by up to 22.9% on unseen tasks, while using 30% fewer tokens. It feels like a genuinely grounded step toward agents that can safely improve their own workflow in production. #AIAgents# #SelfImprovement#
顯示更多