๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    50.2K ํŒฌ
762 commits. 315 contributors. 104 first-timers. vLLM v0.30.0 is live. ๐ŸŽ‰ Highlights: ๐Ÿค– Hybrid-attention hot paths: Kimi K3 streamlines KDA, AttnRes, and MLA; DeepSeek-V4.1-Flash adds MXFP8 KV and async Engram; Qwen3.8-Flash-Next fuses QSA/PLE and cuts sparse-GQA overhead ๐Ÿ—„๏ธ HiSparse adds a host tier beneath sparse-MLA decode; under GPU pressure, only top-k misses return to a per-request hot buffer ๐Ÿ› ๏ธ Model Runner V2 brings EAGLE3-style drafts to pipeline parallelism and extends adaptive verification to every draft-model speculator through online acceptance estimation (#50514#, #52228#) ๐Ÿ–‹๏ธ Dual-key Gumbel-max watermark generation and detection, with per-request opt-out and speculative-decoding support ๐Ÿ†• New models include GLM-5.3-Flash, K2-Horizon, Cohere Compass, and Bailing V3 VL โšก Fast Start keeps post-quantized, TP-sharded weights in a per-GPU daemon; restarts map them over CUDA IPC with --load-format ipc_cache Thread ๐Ÿ‘‡
๋” ๋ณด๊ธฐ