Run DeepSeek V4.1 Flash NVFP4 locally 🐳
We made an NVFP4 build of DeepSeek's v4.1 flash model with
@NVIDIAAI's ModelOpt recipe, every routed expert re-encoded bit-exactly for Blackwell's FP4 tensor cores and measured on 4x B200
NVFP4 build 👇
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
1/6
顯示更多