I really enjoyed
@willccbb's theory-grounded post on the post-training landscape and where OPD fits:
A simple framing I liked:
-SFT gives dense imitation, but lacks on-policy correction.
-RL gives sparse but verifier-grounded correction.
-OPD gives dense on-policy correction when the teacher signal is reliable and compatible.
Our work pushes this one step further: can we get the benefits of OPD without requiring same-family, token-level, logit-based teachers?
We show the answer is yes. By replacing token-logit matching with chunk-level semantic verification, we enable dense on-policy distillation from black-box teachers, with stabilizers that help prevent collapse.
More details in our paper: