Register and share your invite link to earn from video plays and referrals.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
Joined March 2024
36 Following    48.4K Followers
Six weeks ago, serving DSpark in vLLM meant choosing a draft length for your traffic and living with it. Now you set that length once, and vLLM decides how much of the draft to verify every step. ✨ On DeepSeek-V4-Pro-0813, the first token of a 7-token draft survives verification more than 70% of the time. The last one, less than 10%. 📈 One config, adaptive verification on with num_speculative_tokens 7, holds the Pareto frontier from concurrency 1 to 256 on 8×B300. Long draft at low load, short at high load. On main behind enable_adaptive_verification, for DSpark with a confidence head on Flash Attention or DSV4 attention. More backends and models are in bring-up. Thanks to Lucas Wilkinson (@RedHat_AI) and Ben Chislett (@NVIDIAAI), and to @deepseek_ai for DSpark and the varlen DeepGEMM indexer kernel. 🔗
Show more