Register and share your invite link to earn from video plays and referrals.

Damnang
@damnang2
Semiconductor Engineer in SV
Joined October 2025
967 Following    45.2K Followers
I read LS Securities’ report, “Memory Shortages Even in 2030.” (I love this title name) First, thanks to @PhotonCap for sharing it with me. TL;DR: KV cache offload and HBM optimization do not necessarily mean lower overall memory demand. As AI advances, memory is likely to expand beyond HBM into more layers such as local DRAM, CXL, and storage, while becoming increasingly specialized for different workloads. On the supply side, even as fabs and wafer capacity continue to expand, improving bits per wafer through DRAM scaling is becoming increasingly difficult. At the same time, a meaningful portion of incremental DRAM bits may be allocated to HBM, so wafer growth should not be interpreted directly as conventional DRAM supply growth. On the demand side, memory content per system is becoming more important than CPU or GPU unit growth. Even if server shipments do not increase dramatically, more memory channels, higher-density DIMMs, and the expanding memory footprint of AI workloads can continue to drive memory demand per system higher. This becomes particularly important with Agentic AI. As context windows, KV cache, and agent state grow, the memory problem expands from bandwidth alone to capacity. Not everything can reside in HBM, which naturally leads toward a broader memory hierarchy: HBM → Local DRAM → CXL → Storage This also connects with the custom memory thesis I have discussed for some time. Rather than relying on one general-purpose memory architecture, I expect AI systems to increasingly optimize bandwidth, capacity, and latency for specific workloads. This is why I continue to pay attention to architectures such as Custom HBM, Memory-on-Logic, zHBM, and HBF. And as memory becomes more distributed across these tiers, the next bottleneck increasingly becomes data movement. Accessing a larger memory pool is only useful if data can move between these tiers with sufficient bandwidth, low latency, and low power. This is also why I continue to view CXL, fabric, CPO, and optics as closely connected to the evolution of memory architecture. Memory optimization does not mean less memory. It means memory moves. AI systems are becoming more memory-heavy, memory itself is becoming more tiered and specialized, and the fabrics and optics connecting those memory pools should become increasingly important as well.
Show more