昨天卖了个关子,今天可以公布了,新论文叫 CRAFT。
论文里最喜欢的一点,是它不用提前知道用户会问什么。视频压一次,后面换不同的问题,继续用同一份压缩结果。
这对 AI 眼镜很关键。设备不能等你每问一句,再把整段视频从头算一遍。它得先把看过的内容压下来,需要时再拿出来。
约 8× 压缩是论文里的结果。
接下来,更期待看 CRAFT 在真实设备上能省下多少算力和电量。
Another paper from us, this time on video token compression.
Video VLMs have a token problem.
A single video can produce tens of thousands of visual tokens, making prefill expensive and KV caches large. But aggressively pruning them risks deleting the details needed for temporal reasoning.
CRAFT takes a different approach.
It separates two decisions:
Which tokens should merge?
Global similarity decides.
How should they merge?
Position-aware weights and a channel-wise gate learn how to preserve the information.
The process is recursive and query-agnostic, so the video can be compressed once and reused across multiple questions.
At roughly 8× compression, CRAFT retains 96.8% of the backbone’s average accuracy while substantially reducing prefill and KV-cache costs.
The takeaway: don’t just drop redundant tokens. Learn how to fuse them without losing the details.
顯示更多