가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Jon Durbin
@jon_durbin
Human. Backend dev
가입 December 2012
138 팔로잉 중    7.2K 팬
167k tokens per second training throughput on a single 8x 5090 box with parallax (on a toy 5b 1b active test moe)... That's quite high. Every GPU you add reduces compute for the others so that number goes up > linearly. Exciting times. And look at the beautiful loss curve, beating DDP baseline at matched steps (identical warmup/lr/dataset/etc.) With 5.99 GFLOPS per token for this arch and BF16 denominator that's ~60% MFU equivalent. 3.068 nats so far at ~19b tokens (llama-3 tokenizer), beating the DDP baseline at matched token budget and getting there ~15% faster and 64% cheaper, and remember it's ternary routed experts... Best guess as to nats gap is separate expert muon optimizers with the surrogate feedback+fold normalization etc. Almost ready for the big runs...
더 보기