Register and share your invite link to earn from video plays and referrals.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
Joined March 2024
36 Following    45.4K Followers
vLLM hit new peak bs=1 decode on Kimi-K3: 464 tok/s 🚀 Under a low-entropy reasoning workload, Kimi-K3 + DSpark on vLLM reaches 464 tok/s on batch size 1 with 4×4 GB300. This benchmark is fully reproducible with public image: vllm/vllm-openai:kimi-k3 and @inferact's DSpark draft model linked in the thread. 1/3
Show more