注册并分享邀请链接,可获得视频播放与邀请奖励。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在关注    48.4K 粉丝
Six weeks ago, serving DSpark in vLLM meant choosing a draft length for your traffic and living with it. Now you set that length once, and vLLM decides how much of the draft to verify every step. ✨ On DeepSeek-V4-Pro-0813, the first token of a 7-token draft survives verification more than 70% of the time. The last one, less than 10%. 📈 One config, adaptive verification on with num_speculative_tokens 7, holds the Pareto frontier from concurrency 1 to 256 on 8×B300. Long draft at low load, short at high load. On main behind enable_adaptive_verification, for DSpark with a confidence head on Flash Attention or DSV4 attention. More backends and models are in bring-up. Thanks to Lucas Wilkinson (@RedHat_AI) and Ben Chislett (@NVIDIAAI), and to @deepseek_ai for DSpark and the varlen DeepGEMM indexer kernel. 🔗
显示更多
0
16
249
29
转发到社区