登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
参加 June 2015
180 フォロー中    4.1K ファン
In post-training, OPD can sit between SFT and final RL. RL + OPD with fused advantages is also an option, but requires careful annealing as OPD saturates quickly. Either way, the live RL run could have started from an OPD-trained checkpoint.
もっと見る
Hoping someone can explain to me what's going on here. The report says they trained multiple mixRL teachers, then used MOPD to combine those teachers into a student model. But..the team literally live-streamed their RL run, and then released model + report just one day after it finished. How does this timeline work out? I would assume the streamed RL run is just a final climb after MOPD, but the paper doesn't mention anything about a final climb. Maybe the livestreamed run is only for one of their mixRL teachers, and they did MOPD quickly in the 24 hours before launch?
もっと見る