가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
가입 June 2015
180 팔로잉 중    4.1K 팬
In post-training, OPD can sit between SFT and final RL. RL + OPD with fused advantages is also an option, but requires careful annealing as OPD saturates quickly. Either way, the live RL run could have started from an OPD-trained checkpoint.
더 보기
Hoping someone can explain to me what's going on here. The report says they trained multiple mixRL teachers, then used MOPD to combine those teachers into a student model. But..the team literally live-streamed their RL run, and then released model + report just one day after it finished. How does this timeline work out? I would assume the streamed RL run is just a final climb after MOPD, but the paper doesn't mention anything about a final climb. Maybe the livestreamed run is only for one of their mixRL teachers, and they did MOPD quickly in the 24 hours before launch?
더 보기