Register and share your invite link to earn from video plays and referrals.

Haocheng Xi
@HaochengXiUCB
Second-year PhD in @berkeley_ai. Prev: Yao Class, @Tsinghua_Uni | Efficient Machine Learning & ML sys
502 Following    2.3K Followers
Open-source video generation is now faster than playback without compromising quality. Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality. VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11 seconds on 8× NVIDIA B200 GPUs. Checkpoints + training/inference code + Technical Blog ⬇️ (1/6)
Show more
0
66
1.2K
166
Forward to community
🎬 Long video generation is bottlenecked by KV-cache memory. We fixed it. Presenting QuantVideoGen @ #ICML2026#: ⚡ ~7× smaller KV cache @ INT2 💾 ~85% less memory 🔧 No fine-tuning, no weight changes 📍 Poster: Thu Jul 9, 2026. 5:00 PM -6:45 PM KST · Hall A #4603# My wonderful co-author Xingyang Li will be presenting our papers. Come say hi! 🔗
Show more
New blog post: The Forgetting Wall in Video and World Models Long-horizon video generation is not just limited by compute. It is limited by how much of its own past the model can afford to remember. I wrote about why long videos drift, why KV cache becomes the memory bottleneck, and why compression is a key direction for future video/world models.
Show more
Really exciting to see KV-cache compression getting attention. A similar bottleneck shows up beyond LLMs: for world models and autoregressive long-video generation, KV cache can quickly dominate memory and limit long-horizon consistency. Our recent work, Quant VideoGen, explores training-free 2-bit KV-cache quantization for video diffusion models, achieving up to 7.0× KV memory reduction with <4% latency overhead. Link:
Show more