A 1.6T parameter model does not use all 1.6T for every token. And a model that uses 13B per token can still hold far more than 13B.
The card shows how that works. Here is why it matters.
Models usually get better as you add parameters. But using more parameters takes more work. In a dense model, every token uses all the weights. MoE uses only a few of them at a time. That lets the model grow without making each token cost more.
Look at Mixtral and DeepSeek-V4-Flash. Mixtral used 12.9B parameters per token in 2023. V4-Flash uses 13B in 2026. Almost the same number, but V4-Flash holds about six times as many weights. V4-Pro goes further and uses about three percent of its 1.6T.
So the model grew much larger while the part used for each token barely changed. But all those weights still have to sit somewhere, and how much room they take depends on their precision.
Weights are only part of the memory story. Attention design can shrink the KV cache instead. At one million tokens of context, DeepSeek reports that V4-Flash needs seven percent of the KV cache V3.2 needed.
Memory can grow in one place and shrink in another. More weights to hold. Less cache to keep.
This series keeps coming back to that trade. The KV cache, quantization, and now MoE. Each one is a choice about what to keep in memory and what to move.
Less math does not mean less memory.
Earlier in this series: KV cache, prefill and decode, quantization.
顯示更多