註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Rihard Jarc
@RihardJarc
Researching and investing in tech (AI, cloud, semiconductors, platforms). DMs open. Tweets are only opinions.
加入 January 2016
2.8K 正在關注    81.3K 粉絲
A new inference optimized hardware. They won’t have enough $TSM allocation to make serious market share gains in the near term, but this could shake the semis complex a lot. Strong backers. Looks like a complex build. Can help with the HBM bottleneck.
顯示更多
Introducing Cluster-Scale Memory (CSM) for low latency workloads. Today's AI chips using HBM can’t achieve SRAM-level decode speeds due to memory subsystem and interconnect bottlenecks. SRAM-only chips have lower FLOPs density and memory capacity, sacrificing throughput. You’re forced to make a tradeoff: serve at much slower speeds, or run at low batch sizes and suffer from higher costs. When running large MoE models, token routing across experts requires sending data through a deep memory hierarchy and a networking switch to reach a destination expert. Each memory layer inherently adds latency; thus, the best layer is no layer. We’ve designed a new architecture that creates a shared low-latency memory pool across the entire scale-up domain. We use a proprietary ultra-low-latency, high-bandwidth interconnect to enable dramatically faster memory access across chips. Our HBM/SRAM hybrid design solves both memory capacity and mem2mem latency, enabling high throughput and interactivity simultaneously. CSM improves latency and avoids today's cost, reliability, yield, thermal, and compute tradeoffs of SRAM-only chips, 3D DRAM chips, or optics.
顯示更多
0
12
158
11
轉發到社區