註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Haocheng Xi
@HaochengXiUCB
Second-year PhD in @berkeley_ai. Prev: Yao Class, @Tsinghua_Uni | Efficient Machine Learning & ML sys
加入 August 2024
502 正在關注    2.3K 粉絲
Really exciting to see KV-cache compression getting attention. A similar bottleneck shows up beyond LLMs: for world models and autoregressive long-video generation, KV cache can quickly dominate memory and limit long-horizon consistency. Our recent work, Quant VideoGen, explores training-free 2-bit KV-cache quantization for video diffusion models, achieving up to 7.0× KV memory reduction with <4% latency overhead. Link:
顯示更多
0
16
474
65
轉發到社區