from the vault: on-policy distillation in TRL got 40x faster with three fixes, a generation buffer, batched teacher requests, and binary-encoded logprobs
enough to distill from 100B+ teachers, a 4B student gained 39 AIME25 points learning from a 235B one