注册并分享邀请链接,可获得视频播放与邀请奖励。

SemiAnalysis
@SemiAnalysis_
加入 January 2024
34 正在关注    168K 粉丝
Spec Decoding (DSpark) finally works with Pipeline Parallelism in vLLM!! 🔥 A lot of the GPU poor/proletarian have been asking for this feature for a while now, but it took the recent Kimi K3's massive 2.8T parameters affecting the GPU middle class (B200) for it to be implemented! Before this change, GPU poors that needed to use pipeline parallelism, as their weights didn't fully fit on 1 server, would not be able to take advantage of DSpark. Shoutout to yongqinwang-cmd & Inferact for implementing it in vLLM!
显示更多