๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    45.4K ํŒฌ
vLLM hit new peak bs=1 decode on Kimi-K3: 464 tok/s ๐Ÿš€ Under a low-entropy reasoning workload, Kimi-K3 + DSpark on vLLM reaches 464 tok/s on batch size 1 with 4ร—4 GB300. This benchmark is fully reproducible with public image: vllm/vllm-openai:kimi-k3 and @inferact's DSpark draft model linked in the thread. 1/3
๋” ๋ณด๊ธฐ