Your are making claims that extrapolate the actual impact of this.
1. There is no mega kernel equivalent for GPU
2. Its a best case with 6 tokens accepted from draft model
3. Only valid for low volume of request. Is that a real use case?
vLLM is cooking, yes, but not it’s not as easy as that