注册并分享邀请链接,可获得视频播放与邀请奖励。

Ali Hatamizadeh
@ahatamiz1
LLM Tech Lead & Staff Research Scientist @NVIDIA Co-creator of Gated DeltaNet & Gated DeltaNet-2
加入 June 2015
180 正在关注    4.1K 粉丝
I don't believe ( establishes anything about Muon for RLVR. They run GRPO with Qwen3-1.7B/4B over GSM8K, and report that Muon consistently fails. But as always, the devil is in the details. Just quickly checking their setup in their repo, Muon's per-element update RMS comes out at 5×lr. AdamW's is typically 0.2 to 0.4×lr. So they are effectively running Muon at roughly 20x AdamW's step size at the same LR. And of course they did not sweep the optimal LR for Muon in this recipe. We have seen massive success in both pre-training and post-training experiment for Muon and there's no reason to believe otherwise.
显示更多