GLM-5.3 Flash quant showdown on 2× DGX Spark per model.
Bench v6.7.1 — 76 scenarios × 2 repeats.
EXL3 TR3 4bpw
• TrueScore: 90.9
• Capability: 92.2
• Operational: 87.6
• Median turn: 4.30s
• Long-response effective rate: ~35.4 tok/s
• 53,347 total output tokens
NVFP4
• Raw TrueScore: 78.5
• Capability: 71.6
• Operational: 71.9
• Median turn: 13.42s
• Long-response effective rate: ~26.5 tok/s
• 996,608 total output tokens
EXL3 won instruction following (94.2 vs 48.9), structured output (97.2 vs 12.8), code (98.1 vs 68.5), visuals (92.3 vs 36.3), long context (100 vs 85.3) and robustness (100 vs 86.8).
NVFP4 won agentic work (98.5 vs 88.0), planning (100 vs 91.2) and narrowly won safety (87.8 vs 85.6). Tool use tied at 78.4.
NVFP4 completed every request with zero transport errors, but emitted 18.7× more output and took 3.12× longer per median turn.
Important: this is a deployment comparison, not a pure quant-only test. EXL3 used FP8 KV, CUDA graphs and a configured 1M context ceiling. NVFP4 used Marlin, FP8 E4M3 KV, eager mode and a 262K ceiling.
Verdict: EXL3 TR3 4bpw was the much better all-around deployment. NVFP4 was genuinely strong for planning and agentic workflows, but its verbosity and format behavior need fixing before it can compete on practical quality and speed.