註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在關注    56.9K 粉絲
“Why Does Post-Training Quantization Work?” This paper shows that pretrained LLMs are inherently robust to quantization because layers actively counteract accumulated quantization errors, while high-dimensional LM-head geometry preferentially preserves top-token predictions. This lets 4-bit models stay close to full precision despite substantial hidden-state perturbations.
顯示更多
0
3
191
21
轉發到社區