Register and share your invite link to earn from video plays and referrals.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
Joined May 2026
279 Following    414 Followers
What if letting an agent evolve its own workflow just meant it got better at gaming the test? This paper tackles exactly that overfitting problem. Title: RRSI: Regularized Recursive Self-Improvement of Agent Harnesses URL: โ“ What's a harness? It's everything wrapped around the frozen LLM: prompts, control flow, tooling, memory, context management. The same model can perform very differently depending on this design. โ“ Why does self-improving it overfit? Three failure modes show up: fitting too tightly to the evolve-set benchmark, chasing noise, and letting complexity pile up unchecked. ๐Ÿ’ก How does RRSI fix it? It applies classic ML regularization (L0, L1, L2) to harness evolution: an annealed edit budget limits changes per round, benchmark-specific proposals get screened out at selection, and unproductive components get pruned. ๐Ÿ’ก What's the payoff? Across 8 benchmarks, RRSI beats baselines by up to 22.9% on unseen tasks, while using 30% fewer tokens. It feels like a genuinely grounded step toward agents that can safely improve their own workflow in production. #AIAgents# #SelfImprovement#
Show more