Register and share your invite link to earn from video plays and referrals.

Search results for NVFP4
NVFP4 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including NVFP4
"Full" NVFP4 Qwen3.8 27B: Did you try this? I can't trust the eval. LiveCodeBench results for the BF16 baseline are 10~15 points below what they should be. Same for AIME. Evaluated with a low context length (32K), so in a condition where seeing quantization degradation is unlikely Just wondering whether I should spend one or two days on an RTX Pro 6000 to confirm, but at native context length.
Show more
Qwen3.8-27B NVFP4 variants are very close in accuracy. The real differences are memory and speed. If you're VRAM-limited, minima-ai/mnma_qwen3.8_27b_nvfp4 is a good pick, but it doesn't include MTP for faster inference. NVIDIA's version has MTP, but in my long-context coding tests MTP-4 is only ~2× faster than no MTP, and still ~2.5–3× slower than Unsloth (RTX Pro 6000). A likely reason: NVIDIA quantizes lm_head to NVFP4, while Unsloth keeps it FP8. Since MTP shares the target model's lm_head, this can hurt prediction quality and acceptance rate. So my pick is Unsloth NVFP4. Now testing accuracy on long-horizon agentic coding.
Show more
Sparkbench local battle: Qwen3.8-Flash-Next (NVFP4 TP2, spec off) scored 87.0 GLM-5.3-Flash (NVFP4 TP2, DFlash2) at 85.7 This is a deployment comparison, not a universal model ranking, but Qwen was faster, more reliable, and cleaner here. The most interesting result wasn’t the score: GLM hit four ~261K-token repetition loops, while Qwen’s NEXTN path had its own collapse and had to be disabled Local inference moves fast, but serving-stack correctness matters as much as the model. Full deployment and results:
Show more
Minimax H3 😃Video Gen models + workflows (NVFP4/BF16/FP8/INT8/INT4/GGUFs) community released All listed at one place 👇😲
The official GLM-5.2 NVFP4 from NVIDIA is now available. Curious how it compares to other quantizations.
Nvidia dropped an official DeepSeek-V4.1-Flash NVFP4 build. And this thing is BIG. DeepSeek V4.1 Flash stats 🧠 552B backbone 📚 +196B Engram conditional memory ⚡ only 8B active during prefill 🚀 16B active during decode 👁️ native vision 📖 1 MILLION token context 🧩 384 routed experts across 40 layers 📜 MIT license Nvidia has converted its routed MoE experts to NVFP4 W4A4 specifically for Blackwell GPUs. This is not a compressed giant model down to 4-bit. DeepSeek's experts were already stored in MXFP4, Nvidia instead converts them to its Blackwell-friendly NVFP4 format. And because NVFP4 uses finer scaling, the checkpoint actually gets slightly larger. 💾 Source: ~476 GiB 💾 NVIDIA NVFP4: ~492 GiB 48 safetensor shards. 😳 So why bother? Because NVIDIA is optimizing how those 4-bit experts execute on Blackwell. And impressively, NVIDIA's evaluations show basically no obvious quality collapse from the conversion. For example: 🧠 GPQA Diamond 91.04 → 91.29 💻 SciCode 54.40 → 55.84 🛠️ Terminal-Bench 2.1 81.60 → 82.16 👁️ MMMU-Pro 74.05 → 73.70 Some slightly up. Some slightly down. Essentially benchmark parity. And it already has: ✅ vLLM support ✅ SGLang support ✅ reasoning parser ✅ tool calling ✅ image input ✅ 1M context ⚠️ Nvidia validated it on 4× GB300 GPUs so not a local model (yet). The checkpoint is still ~492 GiB. 🔗 HF: /nvidia/DeepSeek-V4.1-Flash-NVFP4
Show more
🔥 GLM-5.3-Flash 320B-A18B MoE · NVFP4 · 4x DGX Spark TP4 Re-benched warmed + streaming: ⚡️ 63.8 tok/s PEAK on structured output (MTP runs hot) 🔢 ~61 median counting · 53 code · 37 prose 🧠 1.26M-token fp8 KV · 262K ctx · 0.2s TTFT Day-0. 320B-class reasoning on desk hardware 🖥️
Show more
🔥 The wildly popular DGX-Spark-friendly Qwen3.8-27B-NVFP4 now comes naked. The OrcaRouter family has already pulled 450K+ Hugging Face downloads across its FP8, GGUF, MLX, NVFP4 and BF16 builds. And now there’s a 23.4GB NVFP4 version built for NVIDIA Blackwell. 👀 🧠 Qwen3.8-27B dense 🚫 Abliterated / dramatically fewer refusals 👁️ Vision preserved 🛠️ Tool calling preserved 🚀 MTP preserved 📚 262K context 📦 ~23.4GB 📜 Apache 2.0 And this isn't a dumb quant conversion. It mixes precision ... ⚡ Most FFN layers → NVFP4 🧮 Attention + DeltaNet → FP8 🧠 Final FFN layers → FP8 💾 KV cache → FP8 👁️ Vision + embeddings + MTP → BF16 Target hardware 🔥 RTX 5090 🔥 DGX Spark / GB10 🔥 RTX PRO Blackwell 🔥 Multi-GPU RTX 50-series 🔥 B200 / B300 For AMD, Intel or Mac, there are also GGUF/MLX versions available 🔗 HF /orcarouter/Qwen3.8-27B-Uncensored-NVFP4
Show more
LTX 2.5 😵😻Open weights released -model (int8-convrot/ BF16/ NVFP4) -temporal/ spatial upscalers -distilled loras -txt encoders/vae 👇
Benchmarking @NVIDIAAI's Nemotron Puzzle 75B locally on the GX10. NVFP4 via vLLM's OpenAI API, MTP speculative decoding, forced 1,500-token generations. 🏃‍♀️22.75 tok/s in a single session into 88.85 cumulative at 7 sessions. 🧍Baseline without MTP: ~16.3. Scripts + setup:
Show more