Congrats to
@Kimi_Moonshot on the Kimi K3 announcement! 🎉
Grateful for the shoutout and collab. The Kimi team announced they contributed a KDA prefix caching implementation directly to vLLM, to be released alongside the model. 🚀
KDA breaks assumptions behind conventional prefix caching, and this upstream work means the community gets efficient long-context serving from day 0.
vLLM will support Kimi K3 on day-0 release.
Open weights by July 27, 2026.