註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

DAIR.AI
@dair_ai
Democratizing AI research, education, and technologies. Learn about AI Agents for FREE at
加入 July 2017
1 正在關注    132.8K 粉絲
Finally a good paper testing whether memory-based self-improving agents actually improve. The re-evaluation adds two things prior work skipped, multiple runs to measure variance and randomly shuffled task orders. Both hurt. Agent evaluation is already noisy on multi-step tasks, and stacking a self-improvement loop amplifies that noise. Default task orderings impose an implicit curriculum that much of the reported gain was riding on. Adding detailed rubrics and environment feedback to memory construction recovers part of the drop, and a significant gap remains. Paper: Track more trending AI papers in our academy:
顯示更多
0
11
93
14
轉發到社區