This is quite an elegant way to increase the probability of low-probability, high-reward trajectories.
TailRL changes the credit assignment so that trajectories reaching reward levels that few other rollouts reach get more weight. Instead of only pulling up expected reward, it optimizes the probability of exceeding reward thresholds across the whole distribution, with a clean finite version of the gradient estimator and a nice best
@k interpretation.
It also reduces exactly to MaxRL when rewards are binary. The paper is very well written and has a ton of experiments across different domains.
larger rollout budgets during training incorporate progressively higher order best
@k , while the tail-focused objective increases the probability of sampling low prob but correct trajectories that become especially useful when you can sample many rollouts at inference.