登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

John Schulman
@johnschulman2
@thinkymachines. Interested in reinforcement learning, alignment, birds, jazz music
参加 May 2021
2.1K フォロー中    79.5K ファン
PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise - the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO,
もっと見る