167k tokens per second training throughput on a single 8x 5090 box with parallax (on a toy 5b 1b active test moe)... That's quite high. Every GPU you add reduces compute for the others so that number goes up > linearly. Exciting times. And look at the beautiful loss curve, beating DDP baseline at matched steps (identical warmup/lr/dataset/etc.)
With 5.99 GFLOPS per token for this arch and BF16 denominator that's ~60% MFU equivalent.
3.068 nats so far at ~19b tokens (llama-3 tokenizer), beating the DDP baseline at matched token budget and getting there ~15% faster and 64% cheaper, and remember it's ternary routed experts...
Best guess as to nats gap is separate expert muon optimizers with the surrogate feedback+fold normalization etc.
Almost ready for the big runs...