Register and share your invite link to earn from video plays and referrals.

Kevin Frans
@kvfrans
phd @berkeley_ai prev mit, reflection, openai read my thoughts:
518 Following    4.3K Followers
Here's a fun one with @preston_fu et al -- how should we reward RL policies in a way that scales to longer and more difficult tasks? Our answer lies in-between RL and imitation learning, and provides a simple way to assign dense per-token credit to long trajectories.
Show more