🚀New blog: Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache
4-bit KV cache is here! NVFP4 KV in SGLang packs ~1.78× more context into GPU memory and speeds up long-context decode by up to 78%.
Together with
@Alibaba_Qwen and
@nvidia, we brought NVFP4 KV cache to Blackwell:
- NVFP4 stores KV in just ~56% of FP8's footprint per token
-️ Decode throughput jumps +37% / +58% / +78% at 32K / 160K / 1M context
- Near-lossless accuracy: matches FP8 on GPQA-Diamond & AIME 2025 (Qwen3.5-397B-A17B)
- Higher cache hit rate keeps AgentX throughput scaling where FP8 drops off
The whole recipe combines NVFP4 two-level scaling, in-kernel dequantization on decode, and paged KV cache, and you can enable it in SGLang with a single flag --kv-cache-dtype nvfp4