๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    45.4K ํŒฌ
Congrats to @Kimi_Moonshot on the Kimi K3 announcement! ๐ŸŽ‰ Grateful for the shoutout and collab. The Kimi team announced they contributed a KDA prefix caching implementation directly to vLLM, to be released alongside the model. ๐Ÿš€ KDA breaks assumptions behind conventional prefix caching, and this upstream work means the community gets efficient long-context serving from day 0. vLLM will support Kimi K3 on day-0 release. Open weights by July 27, 2026.
๋” ๋ณด๊ธฐ