Register and share your invite link to earn from video plays and referrals.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
Joined July 2023
549 Following    11.2K Followers
Nvidia dropped an official DeepSeek-V4.1-Flash NVFP4 build. And this thing is BIG. DeepSeek V4.1 Flash stats 🧠 552B backbone 📚 +196B Engram conditional memory ⚡ only 8B active during prefill 🚀 16B active during decode đŸ‘ī¸ native vision 📖 1 MILLION token context 🧩 384 routed experts across 40 layers 📜 MIT license Nvidia has converted its routed MoE experts to NVFP4 W4A4 specifically for Blackwell GPUs. This is not a compressed giant model down to 4-bit. DeepSeek's experts were already stored in MXFP4, Nvidia instead converts them to its Blackwell-friendly NVFP4 format. And because NVFP4 uses finer scaling, the checkpoint actually gets slightly larger. 💾 Source: ~476 GiB 💾 NVIDIA NVFP4: ~492 GiB 48 safetensor shards. đŸ˜ŗ So why bother? Because NVIDIA is optimizing how those 4-bit experts execute on Blackwell. And impressively, NVIDIA's evaluations show basically no obvious quality collapse from the conversion. For example: 🧠 GPQA Diamond 91.04 → 91.29 đŸ’ģ SciCode 54.40 → 55.84 đŸ› ī¸ Terminal-Bench 2.1 81.60 → 82.16 đŸ‘ī¸ MMMU-Pro 74.05 → 73.70 Some slightly up. Some slightly down. Essentially benchmark parity. And it already has: ✅ vLLM support ✅ SGLang support ✅ reasoning parser ✅ tool calling ✅ image input ✅ 1M context âš ī¸ Nvidia validated it on 4× GB300 GPUs so not a local model (yet). The checkpoint is still ~492 GiB. 🔗 HF: /nvidia/DeepSeek-V4.1-Flash-NVFP4
Show more