註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Kevin Frans
@kvfrans
phd @berkeley_ai prev mit, reflection, openai read my thoughts:
加入 August 2013
518 正在關注    4.3K 粉絲
Here's a fun one with @preston_fu et al -- how should we reward RL policies in a way that scales to longer and more difficult tasks? Our answer lies in-between RL and imitation learning, and provides a simple way to assign dense per-token credit to long trajectories.
顯示更多