가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

alphaXiv
@askalphaxiv
High fidelity research
가입 November 2023
101 팔로잉 중    56.9K 팬
“Why Does Post-Training Quantization Work?” This paper shows that pretrained LLMs are inherently robust to quantization because layers actively counteract accumulated quantization errors, while high-dimensional LM-head geometry preferentially preserves top-token predictions. This lets 4-bit models stay close to full precision despite substantial hidden-state perturbations.
더 보기