註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Jon Durbin
@jon_durbin
Human. Backend dev
加入 December 2012
138 正在關注    7.2K 粉絲
167k tokens per second training throughput on a single 8x 5090 box with parallax (on a toy 5b 1b active test moe)... That's quite high. Every GPU you add reduces compute for the others so that number goes up > linearly. Exciting times. And look at the beautiful loss curve, beating DDP baseline at matched steps (identical warmup/lr/dataset/etc.) With 5.99 GFLOPS per token for this arch and BF16 denominator that's ~60% MFU equivalent. 3.068 nats so far at ~19b tokens (llama-3 tokenizer), beating the DDP baseline at matched token budget and getting there ~15% faster and 64% cheaper, and remember it's ternary routed experts... Best guess as to nats gap is separate expert muon optimizers with the surrogate feedback+fold normalization etc. Almost ready for the big runs...
顯示更多
0
17
145
26
轉發到社區