Is RL optimizing the right objective? 🤔
Should we maximize mean reward? Best-of-k? Which k?
Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲
We introduce Tail-Likelihood Reinforcement Learning (TailRL).
It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs.
🧵 1/n