注册并分享邀请链接,可获得视频播放与邀请奖励。

EverMind
@evermind
Long-term agent memory | self-improving | agent harness
加入 November 2025
21 正在关注    4.3K 粉丝
Our new paper. Self-evolving agents have a measurement problem. The hard part isn’t generating changes to an agent harness. It’s knowing which changes genuinely helped rather than got lucky, overfit, or never activated at all. LLMs diagnose failures and propose changes across prompts, knowledge, runtime, tools, and configuration. But deterministic code owns the credit: validity checks, activation checks, paired significance testing, and a sealed test. Across 7 domains, the evolved harnesses gained +9.0 to +15.5pp on the 6 statistically credited sealed tests, retaining 86–147% of the training gain. What transfers is not one universal harness. It’s the diagnose-and-credit loop.
显示更多