註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    46.7K 粉絲
Congrats to @Kimi_Moonshot on the Kimi K3 announcement! 🎉 Grateful for the shoutout and collab. The Kimi team announced they contributed a KDA prefix caching implementation directly to vLLM, to be released alongside the model. 🚀 KDA breaks assumptions behind conventional prefix caching, and this upstream work means the community gets efficient long-context serving from day 0. vLLM will support Kimi K3 on day-0 release. Open weights by July 27, 2026.
顯示更多
0
6
384
33
轉發到社區