注册并分享邀请链接,可获得视频播放与邀请奖励。

Wall St Engine
@wallstengine
Fast, accurate, consistent stock market news, earnings highlights & more. By Brillinsight. Not financial advice.
加入 March 2022
973 正在关注    196.5K 粉丝
DeepSeek just launched V4.1-Flash, cutting KV-cache HBM requirements by roughly 75% and persistent SSD storage by 87.5% compared with V4-Flash. Its global KV cache is down to 890 bytes per token, versus 3,514 for V4-Flash and 48,068 for V3.2. That is roughly 54x smaller than last December’s model. V4.1-Flash has 552B total parameters, but its new Causal Encoder–Decoder architecture activates just 8B when processing inputs and 16B when generating outputs. It also adds native image understanding, alongside new pretraining methods and larger-scale reinforcement learning. The memory savings matter because AI agents repeatedly reuse context from conversations, documents, code and tool results. The KV cache preserves earlier calculations so the model doesn’t have to redo all that work. At 890 bytes per token, a million tokens would occupy about 890MB of global KV cache. That excludes model weights and other memory needed to run the system. The SSD reduction comes from changing both what gets stored and where. Previously, DeepSeek kept global attention data and short-lived sliding-window attention data in its persistent SSD cache. Both typically stayed there for more than 72 hours, even though the sliding-window data was mainly useful for minutes during an active session. V4.1 moves that short-lived data into a shared pool using 10% of each server’s regular memory. Entries expire after minutes, while global context remains in persistent storage for at least 72 hours. When short-lived data is missing, a new technique called “bounded replay” reconstructs it using a limited window of tokens instead of the much larger computation previously required. Removing that data from SSDs, combined with compressing the remaining global cache, produces the reported 8x storage reduction. Reported benchmark scores include 63.9 on Humanity’s Last Exam with tools, 31.2 on Terminal-Bench 4.0 and 54.8 on Automation-Bench. V4.1-Flash is live through the API as deepseek-flash, with off-peak pricing at half the peak rate. DeepSeek is also replacing V4-Pro with V4.1-Flash starting Sept. 14, ahead of a future V4.1-Pro release. Note: These are reductions in cache requirements, not a 75% cut in total GPU memory or an 87.5% cut in all storage. They make retaining and reusing context substantially cheaper without eliminating the other costs of running an agent.
显示更多