Register and share your invite link to earn from video plays and referrals.

Yuhang Zhou
@YuhangZhou2
Research Scientist @Meta | Phd @ClipUMD
99 Following    190 Followers
Excited to share that SandMLE has been accepted by #COLM2026#! We introduced a multi-agent framework that generates diverse synthetic MLE environments to enable the large-scale on-policy RL training. See you then in SF!
Show more
OmniOPD addresses the key bottleneck of black-box teacher distillation by eliminating the need for teacher logits, while its chunk-level supervision provides a more stable gradient signal.