Register and share your invite link to earn from video plays and referrals.

Arthur Zucker
@art_zucker
Head of transformers @huggingface 🤗
Joined October 2021
679 Following    8.4K Followers
Your are making claims that extrapolate the actual impact of this. 1. There is no mega kernel equivalent for GPU 2. Its a best case with 6 tokens accepted from draft model 3. Only valid for low volume of request. Is that a real use case? vLLM is cooking, yes, but not it’s not as easy as that
Show more