註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

steve hsu
@hsu_steve
Physicist, AI Founder, Manifold Podcast
加入 June 2010
0 正在關注    46.8K 粉絲
RSI WATCH Self-modification of agent policy via "dreaming" = exploration of previous histories using a replay simulator. Everything we know about our own brains via introspection can be tried for AI improvement... "the agent can dream over many alternative exploration strategies before redeploying the improved policy online" Meta-Layer RSI Loop (Dream-RSI): ... a recursive self-improvement loop that continuously collects discovery histories through online exploration, constructs replay simulators from history to refine meta-exploration strategies via dreaming, and redeploys the upgraded policy online Empirical Validation: We conduct experiments to demonstrate that Dream-RSI improves both discovery effectiveness and efficiency in several settings.
顯示更多
0
19
186
21
轉發到社區