注册并分享邀请链接,可获得视频播放与邀请奖励。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在关注    50.2K 粉丝
🚀 vLLM's Humming backend can run Chord, @novita_labs' open-source W4A16 MoE kernels for Kimi K2.x. Kernel gains reach 1.33x on H200 TP8 vs tuned public Humming and 2.15x on B300 EP8 decode vs its untuned default. The indexed path works on compatible vLLM revisions; grouped integration is WIP. Great to see the kernels and benchmarks open-sourced! Details:
显示更多
0
14
92
10
转发到社区