註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
加入 June 2015
180 正在關注    4.1K 粉絲
Personal experience: Muon is very strong for post-training (SFT+RLVR). p.s: always read the details and check code when a paper claims "x did not work". The devil is always in the details.
I don't believe ( establishes anything about Muon for RLVR. They run GRPO with Qwen3-1.7B/4B over GSM8K, and report that Muon consistently fails. But as always, the devil is in the details. Just quickly checking their setup in their repo, Muon's per-element update RMS comes out at 5×lr. AdamW's is typically 0.2 to 0.4×lr. So they are effectively running Muon at roughly 20x AdamW's step size at the same LR. And of course they did not sweep the optimal LR for Muon in this recipe. We have seen massive success in both pre-training and post-training experiment for Muon and there's no reason to believe otherwise.
顯示更多
0
3
151
6
轉發到社區