注册并分享邀请链接,可获得视频播放与邀请奖励。

Kevin Frans
@kvfrans
phd @berkeley_ai prev mit, reflection, openai read my thoughts:
加入 August 2013
518 正在关注    4.3K 粉丝
Here's a fun one with @preston_fu et al -- how should we reward RL policies in a way that scales to longer and more difficult tasks? Our answer lies in-between RL and imitation learning, and provides a simple way to assign dense per-token credit to long trajectories.
显示更多