注册并分享邀请链接,可获得视频播放与邀请奖励。

Shrinivas Ramasubramanian
@stablegradients
加入 September 2021
1.8K 正在关注    404 粉丝
Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
显示更多
0
10
506
83
转发到社区