註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Shrinivas Ramasubramanian
@stablegradients
加入 September 2021
1.8K 正在關注    404 粉絲
Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
顯示更多
0
10
506
83
轉發到社區