가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Kevin Frans
@kvfrans
phd @berkeley_ai prev mit, reflection, openai read my thoughts:
가입 August 2013
518 팔로잉 중    4.3K 팬
Here's a fun one with @preston_fu et al -- how should we reward RL policies in a way that scales to longer and more difficult tasks? Our answer lies in-between RL and imitation learning, and provides a simple way to assign dense per-token credit to long trajectories.
더 보기