註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Wësche
@WescheNex1q
Day time artist and night time AI enthusiast. Building & benchmarking frontier LLMs on 4x DGX Spark clusters + Mac. Creator of Vesica Studio. Houston
加入 January 2013
637 正在關注    2.4K 粉絲
GLM-5.3 Flash quant showdown on 2× DGX Spark per model. Bench v6.7.1 — 76 scenarios × 2 repeats. EXL3 TR3 4bpw • TrueScore: 90.9 • Capability: 92.2 • Operational: 87.6 • Median turn: 4.30s • Long-response effective rate: ~35.4 tok/s • 53,347 total output tokens NVFP4 • Raw TrueScore: 78.5 • Capability: 71.6 • Operational: 71.9 • Median turn: 13.42s • Long-response effective rate: ~26.5 tok/s • 996,608 total output tokens EXL3 won instruction following (94.2 vs 48.9), structured output (97.2 vs 12.8), code (98.1 vs 68.5), visuals (92.3 vs 36.3), long context (100 vs 85.3) and robustness (100 vs 86.8). NVFP4 won agentic work (98.5 vs 88.0), planning (100 vs 91.2) and narrowly won safety (87.8 vs 85.6). Tool use tied at 78.4. NVFP4 completed every request with zero transport errors, but emitted 18.7× more output and took 3.12× longer per median turn. Important: this is a deployment comparison, not a pure quant-only test. EXL3 used FP8 KV, CUDA graphs and a configured 1M context ceiling. NVFP4 used Marlin, FP8 E4M3 KV, eager mode and a 262K ceiling. Verdict: EXL3 TR3 4bpw was the much better all-around deployment. NVFP4 was genuinely strong for planning and agentic workflows, but its verbosity and format behavior need fixing before it can compete on practical quality and speed.
顯示更多