Register and share your invite link to earn from video plays and referrals.

GDP
@bookwormengr
AI model & hardware co-design, Inference economics Safe super intelligence for all All views strictly personal
Joined July 2010
12.3K Following    18.5K Followers
Hardware design distillation! ------------------------------- SuperPods come into fashion in the west, after having been invented in the east! Chinese AI hardware makers like Huawei, Alibaba and others started on this approach that SemiAnalysis greatly documents below - low HBM for each NPU and large scale up domains to compensate for that - at least 2 years ahead of the west. That is why everyone seem to be building SuperPods in China! This allows them to live without much lower HBM. Ox Alpha could serve 100T tokens per day on one such cluster. If you follow this SemiAnalysis report you will notice this design pattern is also emerging among western AI hardware makers. Leading SuperPoD example in China is from Huawei - Huawei 950DT based SuperPoD has 8K plus NPUs in scale up domain. - If you strictly define scale up as domain over which direct memory Load-Store is is possible, even then the scale up domain has 1024 NPUs. - WideEP approach allows deploying humongous models on these super pod. - NPUs have near uniform 3 micro second latency to one another and they can talk any-any with uniform bandwidth. They use optical connectivity, presumably NPO. - Their HiBL and HiZQ memories have 1.6TB/s and 4TB/s bandwidth. And even with that they can serve large models at high interactivity. You can read more about this in this article.
Show more
Long Live the Short King: Why 4-hi HBM Wins Same Bandwidth, Fewer Dies: How 4-hi HBM Cuts Inference Costs and Makes Scarce DRAM Go Further