가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

EverMind
@evermind
Long-term agent memory | self-improving | agent harness
가입 November 2025
21 팔로잉 중    4.3K
Our new paper. Self-evolving agents have a measurement problem. The hard part isn’t generating changes to an agent harness. It’s knowing which changes genuinely helped rather than got lucky, overfit, or never activated at all. LLMs diagnose failures and propose changes across prompts, knowledge, runtime, tools, and configuration. But deterministic code owns the credit: validity checks, activation checks, paired significance testing, and a sealed test. Across 7 domains, the evolved harnesses gained +9.0 to +15.5pp on the 6 statistically credited sealed tests, retaining 86–147% of the training gain. What transfers is not one universal harness. It’s the diagnose-and-credit loop.
더 보기