登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

HongyeJin
@serendip410
Bernese Bear lover. LLM trainer. Philomath @TAMU,@PKU1898
参加 November 2019
249 フォロー中    355 ファン
🚙New blog on saving experts from an untold, silent failure mode in ultra-sparse MoEs : While pursuing ultra-sparse MoE training (top-8/768), we found that lower-layer experts can become functionally useless while training/validation loss and load balance remain normal. All experts in these layers have much smaller weight norms than those in healthy layers. We call this silent expert death[Figure 1]. And it wasn’t just us. Auditing public MoEs, we found similar early-layer collapse signatures. For example, MiMo-v2.5-Pro and Qwen3.5-397B-A17B [Figure 2]. Benchmarks can hide unused capacity. Why? This is not merely a load-balancing failure. We found the learning signal itself disappears. Shallow-expert momentum (the accumulated gradients), as well as the square root of second moment, fell by up to ~3 orders of magnitude while deeper layers stayed stable[Figure 1]. By the time norms collapse, the model has already largely stopped relying on the lower pathway. We introduce LM Loss as Auxiliary Loss (LLAL): temporarily connect an early MoE layer to the LM head during the first few thouands steps then remove it. On 60B /180B models, this simple strategy rescued the lower layers—and the gain persisted after removal. It improved nearly every aspect of expert learning: • Better validation loss [Figure3] • No silent expert collapse [Figure 4] • Better downstream metrics [Figure 4] • Markedly better load balance Check out our blog for more details! A bunch of experiments inside—enjoy the 40-minute read :)
もっと見る