NVIDIA $NVDA just introduced NVHBM IP, promising up to 25% more compute die area, 30% more memory bandwidth and 15% lower HBM power than standard HBM4E.
Here’s our take on how it works and what it could mean for chipmakers and memory vendors.
1. A custom die-to-die PHY
NVHBM replaces the conventional HBM interface with a custom die-to-die PHY.
One possible implementation is to use fewer I/Os or lanes while running each lane at a much higher data rate. Fewer lanes would shrink the PHY block, while higher per-lane speeds would offset the lower lane count and allow total bandwidth to rise by 30%.
Total bandwidth is simply the number of lanes multiplied by the data rate per lane.
2. Moving the memory controller off the XPU
NVHBM moves the memory controller from the XPU into the HBM base die.
The first benefit is the additional compute die area. The second is that some memory-management traffic could be handled inside the HBM stack instead of moving back and forth between the HBM and XPU, reducing interface traffic and the power used to move it.
The trade-off is that the memory controller takes up space on the base die. A smaller custom PHY could offset part of that penalty.
As future HBM base dies move from DRAM-class processes to advanced logic nodes, the PHY should shrink further, reducing the area occupied by the memory controller and interface.
The other trade-off is heat. Moving more logic onto the base die raises its power density, making HBM thermal management even more important and strengthening the case for advanced logic processes.
3. Where the power savings may come from
NVIDIA has not explained exactly how NVHBM cuts HBM power by 15%.
The most obvious source is reduced traffic between the HBM stack and XPU after moving the memory controller onto the base die.
We also suspect the custom PHY has been optimized for efficiency. Fewer lanes running at higher speeds do not automatically reduce power. A shift to advanced logic nodes for the base die could provide another source of savings.
There is also a business case for NVIDIA.
The obvious benefit is additional IP licensing revenue, but the bigger goals may be:
1. Shortening the development cycle for each new HBM generation by reducing the amount of upfront co-design required for custom HBM.
2. Lowering switching costs. A more standardized design would make customers less dependent on a single memory supplier and allow them to qualify more sources.
3. Reducing the premium charged for fully customized HBM.
If NVHBM becomes a widely adopted standard, the likely implications are:
1. NVIDIA and custom XPU developers benefit from better performance and lower costs.
2. Advanced logic processes become more important for HBM base dies, benefiting TSMC $TSM, Samsung Electronics (005930 KS) and Intel $INTC.
3. Samsung Electronics (005930 KS), SK hynix $SKHY and Micron $MU could lose some of the switching-cost advantage and pricing leverage that come with proprietary custom HBM designs.
显示更多