็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅ‚ๅŠ  July 2023
549 ใƒ•ใ‚ฉใƒญใƒผไธญ    11.2K ใƒ•ใ‚กใƒณ
Nvidia dropped an official DeepSeek-V4.1-Flash NVFP4 build. And this thing is BIG. DeepSeek V4.1 Flash stats ๐Ÿง  552B backbone ๐Ÿ“š +196B Engram conditional memory โšก only 8B active during prefill ๐Ÿš€ 16B active during decode ๐Ÿ‘๏ธ native vision ๐Ÿ“– 1 MILLION token context ๐Ÿงฉ 384 routed experts across 40 layers ๐Ÿ“œ MIT license Nvidia has converted its routed MoE experts to NVFP4 W4A4 specifically for Blackwell GPUs. This is not a compressed giant model down to 4-bit. DeepSeek's experts were already stored in MXFP4, Nvidia instead converts them to its Blackwell-friendly NVFP4 format. And because NVFP4 uses finer scaling, the checkpoint actually gets slightly larger. ๐Ÿ’พ Source: ~476 GiB ๐Ÿ’พ NVIDIA NVFP4: ~492 GiB 48 safetensor shards. ๐Ÿ˜ณ So why bother? Because NVIDIA is optimizing how those 4-bit experts execute on Blackwell. And impressively, NVIDIA's evaluations show basically no obvious quality collapse from the conversion. For example: ๐Ÿง  GPQA Diamond 91.04 โ†’ 91.29 ๐Ÿ’ป SciCode 54.40 โ†’ 55.84 ๐Ÿ› ๏ธ Terminal-Bench 2.1 81.60 โ†’ 82.16 ๐Ÿ‘๏ธ MMMU-Pro 74.05 โ†’ 73.70 Some slightly up. Some slightly down. Essentially benchmark parity. And it already has: โœ… vLLM support โœ… SGLang support โœ… reasoning parser โœ… tool calling โœ… image input โœ… 1M context โš ๏ธ Nvidia validated it on 4ร— GB300 GPUs so not a local model (yet). The checkpoint is still ~492 GiB. ๐Ÿ”— HF: /nvidia/DeepSeek-V4.1-Flash-NVFP4
ใ‚‚ใฃใจ่ฆ‹ใ‚‹