Evidence-first AI semiconductor analysis: what new silicon claims prove, where bottlenecks move, and whether gains survive at system and economic scale.
HBF is not cheaper HBM.
At Hot Chips, Oxmiq modeled an all-HBF versus all-HBM 72-GPU rack serving Kimi-K2 (1T parameters, FP4; 1M tokens in, 1K out). At equal modeled rack cost and power, HBF provided ~14× the capacity (294.9 vs 20.7 TB), but only ~0.6× the aggregate bandwidth (922 vs 1,584 TB/s). Peak HBF bandwidth requires 64 KB accesses.
These are Oxmiq’s simulation results, not measured silicon. My take: HBF is a software-managed capacity tier for low-batch MoE expert weights and sparse-attention KV cache, not a bandwidth substitute for HBM.
Source: