Register and share your invite link to earn from video plays and referrals.

Misha Goin
@mgoin_
maintainer @vllm_project, inference perf @RedHat_AI (acq @neuralmagic)
484 Following    2.8K Followers
130,093 tok/s *PER GPU* for DeepSeek V4 Pro on real agentic traces achieved by vLLM!! If you call @vllm_project slow just admit your skill issue lol
I’ll be speaking at the vLLM Conference next week, come say hi in SF
Come get your optimized DSpark for Kimi-K3! We applied our best practices in Speculators to train a very fast SWA drafter, particularly performant on long-context agentic workloads with @vllm_project
GLM 5.2 DSpark update! The full Speculators training run is well underway and we have the epoch-1 checkpoint ready for your GPUs using vLLM nightly: This improves upon the speedup from the preview checkpoint by another 1.5-2x. Stay tuned for more!
Show more
GLM 5.2 DSpark preview is here! ✨ This is the first DSpark speculator for a non-DeepSeek frontier model, trained with Speculators and running on vLLM nightly for ~1.5× faster decode for GLM-5.2-FP8 on 4×B300. Stronger checkpoints to come!
Show more
@MainzOnX paged attention the idea is alive and well :) it's in literally every attention backend now! paged attention the circa 2023 .cu file has been unused for a while so we should stop building it
Today I deleted PagedAttention from vLLM