登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Shrinivas Ramasubramanian
@stablegradients
参加 September 2021
1.8K フォロー中    404 ファン
Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
もっと見る