가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
가입 November 2016
490 팔로잉 중    1.4K 팬
I really enjoyed @willccbb's theory-grounded post on the post-training landscape and where OPD fits: A simple framing I liked: -SFT gives dense imitation, but lacks on-policy correction. -RL gives sparse but verifier-grounded correction. -OPD gives dense on-policy correction when the teacher signal is reliable and compatible. Our work pushes this one step further: can we get the benefits of OPD without requiring same-family, token-level, logit-based teachers? We show the answer is yes. By replacing token-logit matching with chunk-level semantic verification, we enable dense on-policy distillation from black-box teachers, with stabilizers that help prevent collapse. More details in our paper:
더 보기
OmniOPD addresses the key bottleneck of black-box teacher distillation by eliminating the need for teacher logits, while its chunk-level supervision provides a more stable gradient signal.