Main result. At xhigh every quant landed between 88.0 and 90.0% pass
@1 on the full suite:
AWQ INT4: 90.0% NVFP4: 89.3% GGUF Q4_K_M: 89.3% FP8: 88.7% NInfer: 88.0%
Yes, the 4 bit quants scored above the FP8 baseline. McNemar says its a statistical tie, first and last place differ by three tasks out of 150.
I started this run to show you how quantization eats quality. There is nothing to show. The gap between quants is smaller than the gap between reasoning presets.