註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Arthur Zucker
@art_zucker
Head of transformers @huggingface 🤗
加入 October 2021
679 正在關注    8.4K 粉絲
Your are making claims that extrapolate the actual impact of this. 1. There is no mega kernel equivalent for GPU 2. Its a best case with 6 tokens accepted from draft model 3. Only valid for low volume of request. Is that a real use case? vLLM is cooking, yes, but not it’s not as easy as that
顯示更多