注册并分享邀请链接,可获得视频播放与邀请奖励。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在关注    50.2K 粉丝
Qwen3.8-2.4T on vLLM: a Pareto frontier spanning 5K total tokens/s/GPU at high throughput and 180 output tokens/s/user at low latency, across tuned PD configurations on @nvidia GB300 NVL72. Workload: 8K input / 1K output. Drawing on lessons from trial and error, we walk through the tuning decisions step by step: budget KV cache, benchmark prefill and decode separately, then choose topologies and MTP settings for each serving target. Deployment configs are included so you can reproduce the results. Great work from the @NVIDIAAI contributors and the vLLM community! Explore the frontier:
显示更多