가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Russ Salakhutdinov
@rsalakhu
CSO @ Sooth Labs, Professor @ CMU, President Elect ICML Board, Ex-VP of Research @ Meta (Multimodal LLMs, AI Agents), ex-Director of AI at @Apple
가입 January 2015
201 팔로잉 중    128.8K 팬
Check out our new work on Tail-Likelihood Reinforcement Learning (TailRL), extending maximum-likelihood RL from binary to continuous rewards. Rather than optimizing only mean reward, TailRL maximizes the expected log of upper-tail probabilities, naturally placing more weight on rare, high-reward rollouts. Its gradient can also be interpreted as a mixture of Best-of-(k) gradients. TailRL requires only a simple modification to the advantage function, making it easy to integrate into existing RL pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL effectively exploits rare high-reward samples and scales better with increased inference-time sampling. Check out a detailed thread by @stablegradients.
더 보기
Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
더 보기