登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

SemiAnalysis
@SemiAnalysis_
参加 January 2024
35 フォロー中    167.7K ファン
Spec Decoding (DSpark) finally works with Pipeline Parallelism in vLLM!! 🔥 A lot of the GPU poor/proletarian have been asking for this feature for a while now, but it took the recent Kimi K3's massive 2.8T parameters affecting the GPU middle class (B200) for it to be implemented! Before this change, GPU poors that needed to use pipeline parallelism, as their weights didn't fully fit on 1 server, would not be able to take advantage of DSpark. Shoutout to yongqinwang-cmd & Inferact for implementing it in vLLM!
もっと見る