註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Silicon Atlas
@Silicon_Atlas
Evidence-first AI semiconductor analysis: what new silicon claims prove, where bottlenecks move, and whether gains survive at system and economic scale.
加入 March 2021
123 正在關注    2.9K 粉絲
HBF is not cheaper HBM. At Hot Chips, Oxmiq modeled an all-HBF versus all-HBM 72-GPU rack serving Kimi-K2 (1T parameters, FP4; 1M tokens in, 1K out). At equal modeled rack cost and power, HBF provided ~14× the capacity (294.9 vs 20.7 TB), but only ~0.6× the aggregate bandwidth (922 vs 1,584 TB/s). Peak HBF bandwidth requires 64 KB accesses. These are Oxmiq’s simulation results, not measured silicon. My take: HBF is a software-managed capacity tier for low-batch MoE expert weights and sparse-attention KV cache, not a bandwidth substitute for HBM. Source:
顯示更多