NVIDIA B300 servers retail around $350k min, crazy power/cooling requirements/etc. (bad maths) per GPU (8) is $43,750, if you can even find them (narrator: you can't) then pay for collocation etc... 9 PFLOPS dense FP4.
DGX Spark: $4700, plug in anywhere, 1 PFLOP (fp4 sparse).
B300 9ish PFLOP dense FP4 super expensive impossible to collocate anywhere etc. etc., $4861.11 per PFLOP.
DGX Spark $4700 plug in any 15 amp outline in any house, $4700/PFLOP (but sparse).
Memory is slower, so I guess use less of it/less manipulation of it...
Yep, math checks out - pretraining on decentralized DGX Spark with native sparse FP4 (checks notes: yes we're using ternary weights which slot perfectly into sparse FP4 compute) should be a thing.
Time to test it out...
10,000 @nvidia B300 GPUs. India's largest AI Factory.
Together AI and @larsentoubro are building the country's biggest GPU cluster, backing open-source inference, fine-tuning, and training at scale for India's AI-native ecosystem.
The Compute Desk Nvidia B300 GPU-hour index is at its all-time high. As margins on training grow thinner, neoclouds are collectively shifting focus to inference. High-priced B300s are generating ROI by lowering the marginal cost per token generated through batch inference.