Six weeks ago, serving DSpark in vLLM meant choosing a draft length for your traffic and living with it. Now you set that length once, and vLLM decides how much of the draft to verify every step. ✨
On DeepSeek-V4-Pro-0813, the first token of a 7-token draft survives verification more than 70% of the time. The last one, less than 10%.
📈 One config, adaptive verification on with num_speculative_tokens 7, holds the Pareto frontier from concurrency 1 to 256 on 8×B300. Long draft at low load, short at high load.
On main behind enable_adaptive_verification, for DSpark with a confidence head on Flash Attention or DSV4 attention. More backends and models are in bring-up. Thanks to Lucas Wilkinson (
@RedHat_AI) and Ben Chislett (
@NVIDIAAI), and to
@deepseek_ai for DSpark and the varlen DeepGEMM indexer kernel.
🔗