가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

GDP
@bookwormengr
AI model & hardware co-design, Inference economics Safe super intelligence for all All views strictly personal
가입 July 2010
12.3K 팔로잉 중    18.5K 팬
Hardware design distillation! ------------------------------- SuperPods come into fashion in the west, after having been invented in the east! Chinese AI hardware makers like Huawei, Alibaba and others started on this approach that SemiAnalysis greatly documents below - low HBM for each NPU and large scale up domains to compensate for that - at least 2 years ahead of the west. That is why everyone seem to be building SuperPods in China! This allows them to live without much lower HBM. Ox Alpha could serve 100T tokens per day on one such cluster. If you follow this SemiAnalysis report you will notice this design pattern is also emerging among western AI hardware makers. Leading SuperPoD example in China is from Huawei - Huawei 950DT based SuperPoD has 8K plus NPUs in scale up domain. - If you strictly define scale up as domain over which direct memory Load-Store is is possible, even then the scale up domain has 1024 NPUs. - WideEP approach allows deploying humongous models on these super pod. - NPUs have near uniform 3 micro second latency to one another and they can talk any-any with uniform bandwidth. They use optical connectivity, presumably NPO. - Their HiBL and HiZQ memories have 1.6TB/s and 4TB/s bandwidth. And even with that they can serve large models at high interactivity. You can read more about this in this article.
더 보기
Long Live the Short King: Why 4-hi HBM Wins Same Bandwidth, Fewer Dies: How 4-hi HBM Cuts Inference Costs and Makes Scarce DRAM Go Further