Register and share your invite link to earn from video plays and referrals.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
Joined May 2026
280 Following    415 Followers
๐Ÿš€ A model that handles a million-token context while shrinking its KV cache to a quarter of the previous size just dropped. DeepSeek's latest. Title: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression URL: ๐Ÿ’ก Overview A 552B multimodal MoE model built for input-heavy, long-running agent workloads, aggressively compressing the KV cache through both architecture and precision optimizations. ๐ŸŽฏ The problem it solves For agents working with long contexts, persistent KV cache storage โ€” not just prefill compute โ€” strains GPU memory, SSD, and I/O bandwidth, becoming the main bottleneck to lowering deployment costs. ๐Ÿ›  Method A Causal Encoder-Decoder structure (first 20 layers as encoder, last 20 as decoder) halves prefill compute. Compressed Sparse Attention 2 reuses cross-layer KV in three modes, a hierarchical sparse indexer narrows search to a candidate pool, and FP4 quantization trims memory further. ๐Ÿ“Š Results Global KV footprint drops to 890 bytes per token, about 1/4 of V4-Flash, with persistent KV around 1/8. Extending context 256x from 4K to 1M only increases decode FLOPs by 1/4. Benchmark gains hold too, e.g. 79.4% on HumanEval. #LLM# #KVCache#
Show more