注册并分享邀请链接,可获得视频播放与邀请奖励。

Arthur Zucker
@art_zucker
Head of transformers @huggingface 🤗
加入 October 2021
679 正在关注    8.4K 粉丝
Your are making claims that extrapolate the actual impact of this. 1. There is no mega kernel equivalent for GPU 2. Its a best case with 6 tokens accepted from draft model 3. Only valid for low volume of request. Is that a real use case? vLLM is cooking, yes, but not it’s not as easy as that
显示更多