vLLM hit new peak bs=1 decode on Kimi-K3: 464 tok/s ๐
Under a low-entropy reasoning workload, Kimi-K3 + DSpark on vLLM reaches 464 tok/s on batch size 1 with 4ร4 GB300.
This benchmark is fully reproducible with public image: vllm/vllm-openai:kimi-k3 and
@inferact's DSpark draft model linked in the thread.
1/3