Qwen3.8-2.4T-A95B, a 2.4-trillion parameter MoE, quantized with no accuracy loss.
MoE layers to NVFP4, attention layers to FP8 block, via LLM Compressor. On GPQA Diamond: 93.1 quantized vs 92.6 for the full-precision base. Full recovery.
Serve on vLLM: