TL;DR: Existing on-policy distillation methods that try to surpass the teacher destabilize training by amplifying noise in output space. A new method, RIDE, instead extrapolates the RL-induced representation change directly in representation space โ and beats the teacher on all four tested model pairs.
Title: The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
URL:
Points
๐ฏ Treats the teacher as a direction, not a destination โ regresses the student toward a target that extrapolates the pre/post-RL hidden-state residual
โ ๏ธ The LM head's anisotropic spectrum attenuates most of that change: the weakest 512 head directions carry 79.8% of hidden-state energy but only 30.0% of output-space energy
๐ Output-space extrapolation's variance grows as (ฮป-1)ยฒ; RIDE's gradient variance is 12.6x lower than ExOPD's
๐ Beats the teacher on all 4 model pairs: 56.38 on R1-Distill-1.5B, 66.07 on Qwen3-4B, outperforming OPRD by 0.97โ4.06 points
โ The output-space baseline ExOPD underperforms both the teacher and OPRD on every pair
๐ฌ Ablations confirm the direction itself matters: random or reversed directions don't deliver the gain that the true RL-induced direction does
It's striking how much the choice of where you measure the difference changes whether "surpassing the teacher" actually works.
#
Distillation# #
ReinforcementLearning#