가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Jon Durbin
@jon_durbin
Human. Backend dev
가입 December 2012
138 팔로잉 중    7.2K 팬
Ternary and sparse FP4 are a match made in heaven. 2x the performance, cost is basically nothing in terms of quality since ternary already has natural sparsity. AND, you can pack those params 1.15 bits per param instead of 1.58. A petaflop... at 200 watts... Almost magic.
더 보기
NVIDIA B300 servers retail around $350k min, crazy power/cooling requirements/etc. (bad maths) per GPU (8) is $43,750, if you can even find them (narrator: you can't) then pay for collocation etc... 9 PFLOPS dense FP4. DGX Spark: $4700, plug in anywhere, 1 PFLOP (fp4 sparse). B300 9ish PFLOP dense FP4 super expensive impossible to collocate anywhere etc. etc., $4861.11 per PFLOP. DGX Spark $4700 plug in any 15 amp outline in any house, $4700/PFLOP (but sparse). Memory is slower, so I guess use less of it/less manipulation of it... Yep, math checks out - pretraining on decentralized DGX Spark with native sparse FP4 (checks notes: yes we're using ternary weights which slot perfectly into sparse FP4 compute) should be a thing. Time to test it out...
더 보기