Register and share your invite link to earn from video plays and referrals.

Search results for V4_1Flash
V4_1Flash community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including V4_1Flash
V4.1-Flash-Vision is a zany model. It has better sight, it's vastly faster, but… 1) it overthinks. Even when it's right in the first 200 words. 2) it's CHEEKY! "Whale-chan, OwO what's dis?" "cool pic, let me think… ah I found the original Codex dir, nvm" wtf whale?
Show more
DeepSeek v4.1 flash made using only Javascript it's a good model sir
DeepSeek v4.1 Flash created this using only code
DeepSeek V4.1 Flash is the model with the biggest gap between its significance and the interest of the evaluator community. No ARC-AGI, no math-arena, no WeirdML… I guess a "0.1 flash" update doesn't sound like big news, plus AA score is middling. Disappointing.
Show more
DeepSeek V4.1 Flash Is an Architecture Reset, Not Just a Cheaper Model Despite having more total and active parameters than its predecessor, V4.1 Flash cuts working KV cache to one quarter and persistent cache storage to one eighth. Zhihu contributor Exhalation explains how it does this by deleting old modules, sharing KV states, and recomputing local context. 1️⃣ DeepSeek removed its own previous ideas 🔹 MTP: external draft models such as DSpark weakened its speculative-decoding value, while the auxiliary loss no longer justified its memory cost. 🔹 Heavily Compressed Attention: its global-summary role was ambiguous and difficult to combine with FP4 storage. 🔹 Dense warmup: V4.1 trains sparse attention from scratch rather than starting with one trillion dense-attention tokens. 2️⃣ Store less, reuse more Non-SWA KV cache moves from FP8 to FP4, while the more sensitive SWA portion remains FP8. DeepSeek no longer persists SWA cache. When a conversation forks from an earlier point, the system rebuilds only a small local window. Post-training simulated this process to limit numerical drift. A modified YOCO design provides the other major saving. Upper layers reuse the same lower-layer source representation, adding only a layer-specific projection. The result is roughly half the KV storage and close to 50% less historical prefill computation in the idealized case. 3️⃣ Sparse attention reuses its search Sparse attention lowers attention cost from O(n²) to O(kn), but finding the top-k tokens can still retain an O(n²) component. V4.1 Flash either reuses an earlier layer’s top-k result or selects a smaller candidate-block pool before re-indexing. This prevents token selection from becoming the bottleneck at long context lengths. 4️⃣ Engram and mHC were streamlined Engram replaces an expensive second-order optimizer state with a Sinkhorn-style update, removes causal convolution, and extends matching from 3-grams to 4-grams. mHC reorders residual mixing across layers, reducing estimated I/O from (4n+4)d to (3n+2)d. ✅ The larger pattern DeepSeek is not merely compressing an existing model. It is willing to discard its own previous components when a cheaper system-level design emerges. V4.1 Flash is less a smaller V4 than a new answer to one question: how much intelligence can be delivered per byte of memory and unit of inference cost? 🔗 Full analysis: #DeepSeek# #DeepSeekV41# #LLMArchitecture# #AIInfra# #KVCache# #SparseAttention#
Show more
DeepSeek-V4.1-Flash is live on @OpenRouter!! pick @wafer_ai as your provider (ss taken 9.12.26)
DeepSeek V4.1 Flash can build a "mechanically accurate pickup truck" in Blender quite well - soon to be run 100% locally 👀 the 🐳 has shipped something really unique here 🫶
Show more
0
52
1.4K
83
Forward to community
DeepSeek v4.1 Flash support is now pushed on DwarfStar "main" branch on GitHub, and this is a YouTube video (in English language) where I test both the SSD streamed and the dual MacBook m5 max 128GB setup during a coding session:
Show more
DeepSeek V4.1 Flash - on pace for the most tokens of any paid model launch in the first 48 hours 😤
DeepSeek V4.1 Flash: 1T tokens in 24 hours on OpenRouter, and is on pace for the biggest 48 hours of any paid model launch at ~2.8T. 90% of those tokens were cache reads, priced by the market at ~$0.006/M, 5x cheaper than comparable models like GLM-5.3 Flash
Show more