🚀 A model that handles a million-token context while shrinking its KV cache to a quarter of the previous size just dropped. DeepSeek's latest.
Title: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
URL:
💡 Overview
A 552B multimodal MoE model built for input-heavy, long-running agent workloads, aggressively compressing the KV cache through both architecture and precision optimizations.
🎯 The problem it solves
For agents working with long contexts, persistent KV cache storage — not just prefill compute — strains GPU memory, SSD, and I/O bandwidth, becoming the main bottleneck to lowering deployment costs.
🛠 Method
A Causal Encoder-Decoder structure (first 20 layers as encoder, last 20 as decoder) halves prefill compute. Compressed Sparse Attention 2 reuses cross-layer KV in three modes, a hierarchical sparse indexer narrows search to a candidate pool, and FP4 quantization trims memory further.
📊 Results
Global KV footprint drops to 890 bytes per token, about 1/4 of V4-Flash, with persistent KV around 1/8. Extending context 256x from 4K to 1M only increases decode FLOPs by 1/4. Benchmark gains hold too, e.g. 79.4% on HumanEval.
#
LLM# #
KVCache#