HBF is not cheaper HBM.
At Hot Chips, Oxmiq modeled an all-HBF versus all-HBM 72-GPU rack serving Kimi-K2 (1T parameters, FP4; 1M tokens in, 1K out). At equal modeled rack cost and power, HBF provided ~14× the capacity (294.9 vs 20.7 TB), but only ~0.6× the aggregate bandwidth (922 vs 1,584 TB/s). Peak HBF bandwidth requires 64 KB accesses.
These are Oxmiq’s simulation results, not measured silicon. My take: HBF is a software-managed capacity tier for low-batch MoE expert weights and sparse-attention KV cache, not a bandwidth substitute for HBM.
Source:
HBF TIME!
SanDisk $SNDK says its HBF architecture can deliver HBM-like read bandwidth with 8–16x more capacity, targeting increasingly memory-heavy AI inference workloads.
The company sees HBF being deployed several ways: replacing some HBM stacks, sitting alongside HBM as higher-capacity read-optimized memory, or using HBM as a cache while HBF stores model weights and KV cache.
SanDisk also sees HBF being used specifically on the decode side for weights and KV cache.
SanDisk $SNDK says HBF could materially reduce GPU requirements for AI inference.
8x capex efficiency
1 HBF GPU ran the same model that required 8 HBM GPUs.
2x GPU efficiency
4 HBF GPUs matched the token output of 8 HBM GPUs.
Based on internal testing.
SanDisk $SNDK HBF simulation delivers the same 12.8 TB/s memory bandwidth as HBM, but with 4TB of capacity per GPU vs. 192GB. For a 960GB Qwen3 workload, HBF needed just 1–4 GPUs vs. 8 with HBM.