Register and share your invite link to earn from video plays and referrals.

SemiAnalysis
@SemiAnalysis_
Joined January 2024
35 Following    167.6K Followers
Spec Decoding (DSpark) finally works with Pipeline Parallelism in vLLM!! 🔥 A lot of the GPU poor/proletarian have been asking for this feature for a while now, but it took the recent Kimi K3's massive 2.8T parameters affecting the GPU middle class (B200) for it to be implemented! Before this change, GPU poors that needed to use pipeline parallelism, as their weights didn't fully fit on 1 server, would not be able to take advantage of DSpark. Shoutout to yongqinwang-cmd & Inferact for implementing it in vLLM!
Show more