登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Russ Salakhutdinov
@rsalakhu
CSO @ Sooth Labs, Professor @ CMU, President Elect ICML Board, Ex-VP of Research @ Meta (Multimodal LLMs, AI Agents), ex-Director of AI at @Apple
参加 January 2015
201 フォロー中    128.8K ファン
Check out our new work on Tail-Likelihood Reinforcement Learning (TailRL), extending maximum-likelihood RL from binary to continuous rewards. Rather than optimizing only mean reward, TailRL maximizes the expected log of upper-tail probabilities, naturally placing more weight on rare, high-reward rollouts. Its gradient can also be interpreted as a mixture of Best-of-(k) gradients. TailRL requires only a simple modification to the advantage function, making it easy to integrate into existing RL pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL effectively exploits rare high-reward samples and scales better with increased inference-time sampling. Check out a detailed thread by @stablegradients.
もっと見る
Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
もっと見る