注册并分享邀请链接,可获得视频播放与邀请奖励。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在关注    50.3K 粉丝
New blog! 🚀 MTP, EAGLE-3, DFlash or DSpark, which speculative decoding method should you actually use? There’s no universal winner. The best choice changes with the model, workload, and speculation depth. We break down how 5 methods work, how to enable and tune them in vLLM, and benchmark them across Gemma, Qwen, Kimi and MiniMax on @AMD Instinct MI300X & MI355X. Deep dive 👇
显示更多
0
17
413
55
转发到社区