登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Kevin Frans
@kvfrans
phd @berkeley_ai prev mit, reflection, openai read my thoughts:
参加 August 2013
518 フォロー中    4.3K ファン
Here's a fun one with @preston_fu et al -- how should we reward RL policies in a way that scales to longer and more difficult tasks? Our answer lies in-between RL and imitation learning, and provides a simple way to assign dense per-token credit to long trajectories.
もっと見る