註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
加入 February 2026
1.9K 正在關注    14.6K 粉絲
On-policy distillation has the same systems bottleneck as RL: rollouts dominate training time on reasoning workloads. Going async fixes throughput but feeds the learner stale-policy data, and what staleness does to OPD specifically was unstudied. The clean finding is that KL direction decides robustness. Teacher-weighted forward KL shrugs off stale rollouts, student-weighted reverse KL breaks under them, and for the reverse-KL case nothing from async RL beats just recomputing the signal under the current student. Finite teacher-score caches then turn the estimator into a bias-variance tradeoff, which is the case for multi-sample Monte Carlo. AsyncOPD: How Stale Can On-Policy Distillation Be? Paper:
顯示更多
0
5
306
44
轉發到社區