註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

John Schulman
@johnschulman2
@thinkymachines. Interested in reinforcement learning, alignment, birds, jazz music
加入 May 2021
2.1K 正在關注    79.5K 粉絲
PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise - the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO,
顯示更多
0
14
1.3K
110
轉發到社區