Check out our new work on Tail-Likelihood Reinforcement Learning (TailRL), extending maximum-likelihood RL from binary to continuous rewards.
Rather than optimizing only mean reward, TailRL maximizes the expected log of upper-tail probabilities, naturally placing more weight on rare, high-reward rollouts. Its gradient can also be interpreted as a mixture of Best-of-(k) gradients.
TailRL requires only a simple modification to the advantage function, making it easy to integrate into existing RL pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL effectively exploits rare high-reward samples and scales better with increased inference-time sampling.
Check out a detailed thread by
@stablegradients.