Register and share your invite link to earn from video plays and referrals.

Jon Durbin
@jon_durbin
Human. Backend dev
Joined December 2012
138 Following    7.2K Followers
Napkin math on bits per byte vocab adjusted etc, this is also beating olmoe 7b 1ba at matched tokens (which is 30% more active params, 40% more total params). Maybe competitive with MobileMoE if we did a complete several trillion token severely overtrained run. Quantile balancing LatentMoE Hybrid GDN-2/SWAX/MSA Role differentiated mixed precision I think we have a winner. More napkin math, I think this 5b model would be able to do around 1200tps inference on 5090s. Kinda want to do the full 9t token run on some 5090s of this tiny one, for funsies. Parallax-5b-instant (or we could bump total params up to 10b for free if we don't change active params, or maybe higher)
Show more
167k tokens per second training throughput on a single 8x 5090 box with parallax (on a toy 5b 1b active test moe)... That's quite high. Every GPU you add reduces compute for the others so that number goes up > linearly. Exciting times. And look at the beautiful loss curve, beating DDP baseline at matched steps (identical warmup/lr/dataset/etc.) With 5.99 GFLOPS per token for this arch and BF16 denominator that's ~60% MFU equivalent. 3.068 nats so far at ~19b tokens (llama-3 tokenizer), beating the DDP baseline at matched token budget and getting there ~15% faster and 64% cheaper, and remember it's ternary routed experts... Best guess as to nats gap is separate expert muon optimizers with the surrogate feedback+fold normalization etc. Almost ready for the big runs...
Show more