ๆณจๅ†Œๅนถๅˆ†ไบซ้‚€่ฏท้“พๆŽฅ๏ผŒๅฏ่Žทๅพ—่ง†้ข‘ๆ’ญๆ”พไธŽ้‚€่ฏทๅฅ–ๅŠฑใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅŠ ๅ…ฅ July 2023
549 ๆญฃๅœจๅ…ณๆณจ    11.2K ็ฒ‰ไธ
Nvidia dropped an official DeepSeek-V4.1-Flash NVFP4 build. And this thing is BIG. DeepSeek V4.1 Flash stats ๐Ÿง  552B backbone ๐Ÿ“š +196B Engram conditional memory โšก only 8B active during prefill ๐Ÿš€ 16B active during decode ๐Ÿ‘๏ธ native vision ๐Ÿ“– 1 MILLION token context ๐Ÿงฉ 384 routed experts across 40 layers ๐Ÿ“œ MIT license Nvidia has converted its routed MoE experts to NVFP4 W4A4 specifically for Blackwell GPUs. This is not a compressed giant model down to 4-bit. DeepSeek's experts were already stored in MXFP4, Nvidia instead converts them to its Blackwell-friendly NVFP4 format. And because NVFP4 uses finer scaling, the checkpoint actually gets slightly larger. ๐Ÿ’พ Source: ~476 GiB ๐Ÿ’พ NVIDIA NVFP4: ~492 GiB 48 safetensor shards. ๐Ÿ˜ณ So why bother? Because NVIDIA is optimizing how those 4-bit experts execute on Blackwell. And impressively, NVIDIA's evaluations show basically no obvious quality collapse from the conversion. For example: ๐Ÿง  GPQA Diamond 91.04 โ†’ 91.29 ๐Ÿ’ป SciCode 54.40 โ†’ 55.84 ๐Ÿ› ๏ธ Terminal-Bench 2.1 81.60 โ†’ 82.16 ๐Ÿ‘๏ธ MMMU-Pro 74.05 โ†’ 73.70 Some slightly up. Some slightly down. Essentially benchmark parity. And it already has: โœ… vLLM support โœ… SGLang support โœ… reasoning parser โœ… tool calling โœ… image input โœ… 1M context โš ๏ธ Nvidia validated it on 4ร— GB300 GPUs so not a local model (yet). The checkpoint is still ~492 GiB. ๐Ÿ”— HF: /nvidia/DeepSeek-V4.1-Flash-NVFP4
ๆ˜พ็คบๆ›ดๅคš