๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    50.2K ํŒฌ
New blog! ๐Ÿš€ MTP, EAGLE-3, DFlash or DSpark, which speculative decoding method should you actually use? Thereโ€™s no universal winner. The best choice changes with the model, workload, and speculation depth. We break down how 5 methods work, how to enable and tune them in vLLM, and benchmark them across Gemma, Qwen, Kimi and MiniMax on @AMD Instinct MI300X & MI355X. Deep dive ๐Ÿ‘‡
๋” ๋ณด๊ธฐ