註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    50.2K 粉絲
Qwen3.8-2.4T on vLLM: a Pareto frontier spanning 5K total tokens/s/GPU at high throughput and 180 output tokens/s/user at low latency, across tuned PD configurations on @nvidia GB300 NVL72. Workload: 8K input / 1K output. Drawing on lessons from trial and error, we walk through the tuning decisions step by step: budget KV cache, benchmark prefill and decode separately, then choose topologies and MTP settings for each serving target. Deployment configs are included so you can reproduce the results. Great work from the @NVIDIAAI contributors and the vLLM community! Explore the frontier:
顯示更多