Register and share your invite link to earn from video plays and referrals.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
Joined May 2026
280 Following    422 Followers
TL;DR: Existing on-policy distillation methods that try to surpass the teacher destabilize training by amplifying noise in output space. A new method, RIDE, instead extrapolates the RL-induced representation change directly in representation space โ€” and beats the teacher on all four tested model pairs. Title: The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation URL: Points ๐ŸŽฏ Treats the teacher as a direction, not a destination โ€” regresses the student toward a target that extrapolates the pre/post-RL hidden-state residual โš ๏ธ The LM head's anisotropic spectrum attenuates most of that change: the weakest 512 head directions carry 79.8% of hidden-state energy but only 30.0% of output-space energy ๐Ÿ“‰ Output-space extrapolation's variance grows as (ฮป-1)ยฒ; RIDE's gradient variance is 12.6x lower than ExOPD's ๐Ÿ“Š Beats the teacher on all 4 model pairs: 56.38 on R1-Distill-1.5B, 66.07 on Qwen3-4B, outperforming OPRD by 0.97โ€“4.06 points โŒ The output-space baseline ExOPD underperforms both the teacher and OPRD on every pair ๐Ÿ”ฌ Ablations confirm the direction itself matters: random or reversed directions don't deliver the gain that the true RL-induced direction does It's striking how much the choice of where you measure the difference changes whether "surpassing the teacher" actually works. #Distillation# #ReinforcementLearning#
Show more