註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
加入 May 2018
2.1K 正在關注    11.7K 粉絲
Lossless quantization has usually meant giving up inference speedup. This paper changes that. SLQ (Statistically-Lossless Quantization) reaches task-lossless compression at 3.3 bits per parameter, and distribution-lossless at 5-6 bpp where the output distribution is practically indistinguishable from the original. 1.7 to 3.6x throughput over BF16 in @vllm_project. Beats FP8 while staying lossless. From Michael Helcig, @_EldarKurtic, and @DAlistarh.
顯示更多
0
4
95
16
轉發到社區