Register and share your invite link to earn from video plays and referrals.

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
Joined November 2016
490 Following    1.4K Followers
I really enjoyed @willccbb's theory-grounded post on the post-training landscape and where OPD fits: A simple framing I liked: -SFT gives dense imitation, but lacks on-policy correction. -RL gives sparse but verifier-grounded correction. -OPD gives dense on-policy correction when the teacher signal is reliable and compatible. Our work pushes this one step further: can we get the benefits of OPD without requiring same-family, token-level, logit-based teachers? We show the answer is yes. By replacing token-logit matching with chunk-level semantic verification, we enable dense on-policy distillation from black-box teachers, with stabilizers that help prevent collapse. More details in our paper:
Show more
OmniOPD addresses the key bottleneck of black-box teacher distillation by eliminating the need for teacher logits, while its chunk-level supervision provides a more stable gradient signal.