Minima AI has released a fully NVFP4-quantized build of Qwen3.8-27B.
Unlike earlier 4-bit quantizations that kept GDN at 8/16-bit, Minima quantized all 496 backbone linear layers to NVFP4 W4A4 — including the GDN layers and their gate projections.
The result: just a 0.52-point average drop across five tasks vs. BF16, while cutting weight memory from 50.13 GiB to 17.53 GiB.
For 32K-token inputs, TTFT also drops from 6.90s to 4.03s.
The interesting part: a 32K-token mechanism study found that quantization error does not keep accumulating in GDN’s recurrent state.
The quantized weights are now open-sourced.
#
Qwen# #
NVFP4#