“Why Does Post-Training Quantization Work?”
This paper shows that pretrained LLMs are inherently robust to quantization because layers actively counteract accumulated quantization errors, while high-dimensional LM-head geometry preferentially preserves top-token predictions.
This lets 4-bit models stay close to full precision despite substantial hidden-state perturbations.
顯示更多