Register and share your invite link to earn from video plays and referrals.

John Schulman
@johnschulman2
@thinkymachines. Interested in reinforcement learning, alignment, birds, jazz music
Joined May 2021
2.1K Following    79.5K Followers
PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise - the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO,
Show more
0
14
1.3K
110
Forward to community