登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
参加 November 2016
490 フォロー中    1.4K ファン
I really enjoyed @willccbb's theory-grounded post on the post-training landscape and where OPD fits: A simple framing I liked: -SFT gives dense imitation, but lacks on-policy correction. -RL gives sparse but verifier-grounded correction. -OPD gives dense on-policy correction when the teacher signal is reliable and compatible. Our work pushes this one step further: can we get the benefits of OPD without requiring same-family, token-level, logit-based teachers? We show the answer is yes. By replacing token-logit matching with chunk-level semantic verification, we enable dense on-policy distillation from black-box teachers, with stabilizers that help prevent collapse. More details in our paper:
もっと見る
OmniOPD addresses the key bottleneck of black-box teacher distillation by eliminating the need for teacher logits, while its chunk-level supervision provides a more stable gradient signal.