注册并分享邀请链接,可获得视频播放与邀请奖励。

Haocheng Xi
@HaochengXiUCB
Second-year PhD in @berkeley_ai. Prev: Yao Class, @Tsinghua_Uni | Efficient Machine Learning & ML sys
加入 August 2024
502 正在关注    2.3K 粉丝
Really exciting to see KV-cache compression getting attention. A similar bottleneck shows up beyond LLMs: for world models and autoregressive long-video generation, KV cache can quickly dominate memory and limit long-horizon consistency. Our recent work, Quant VideoGen, explores training-free 2-bit KV-cache quantization for video diffusion models, achieving up to 7.0× KV memory reduction with <4% latency overhead. Link:
显示更多
0
16
474
65
转发到社区