Excited to share that SandMLE has been accepted by #COLM2026#! We introduced a multi-agent framework that generates diverse synthetic MLE environments to enable the large-scale on-policy RL training. See you then in SF!
OmniOPD addresses the key bottleneck of black-box teacher distillation by eliminating the need for teacher logits, while its chunk-level supervision provides a more stable gradient signal.