Qwen3.8-2.4T on vLLM: a Pareto frontier spanning 5K total tokens/s/GPU at high throughput and 180 output tokens/s/user at low latency, across tuned PD configurations on
@nvidia GB300 NVL72. Workload: 8K input / 1K output.
Drawing on lessons from trial and error, we walk through the tuning decisions step by step: budget KV cache, benchmark prefill and decode separately, then choose topologies and MTP settings for each serving target. Deployment configs are included so you can reproduce the results.
Great work from the
@NVIDIAAI contributors and the vLLM community!
Explore the frontier: