Kimi K3's KDA breaks conventional prefix caching: the recurrent state is cumulative, so per-block snapshots blow up memory. vLLM's fix — decoupling state-block size from prefix-match granularity — enables partial cache hits. Day-0 serving, ready before the Jul 27 weights. JP deep-dive 👇