Qwen3.8-2.4T-A95B, compressed two ways at once: 25% of the experts pruned with REAP, and the rest quantized to NVFP4.
Even with a quarter of the experts gone and 4-bit weights, GPQA Diamond holds at 91.5 vs 92.6 for the full-precision base. ~99% recovery.
Serve on
@vllm_project: