Register and share your invite link to earn from video plays and referrals.

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
Joined June 2015
180 Following    4.1K Followers
In post-training, OPD can sit between SFT and final RL. RL + OPD with fused advantages is also an option, but requires careful annealing as OPD saturates quickly. Either way, the live RL run could have started from an OPD-trained checkpoint.
Show more
Hoping someone can explain to me what's going on here. The report says they trained multiple mixRL teachers, then used MOPD to combine those teachers into a student model. But..the team literally live-streamed their RL run, and then released model + report just one day after it finished. How does this timeline work out? I would assume the streamed RL run is just a final climb after MOPD, but the paper doesn't mention anything about a final climb. Maybe the livestreamed run is only for one of their mixRL teachers, and they did MOPD quickly in the 24 hours before launch?
Show more