Register and share your invite link to earn from video plays and referrals.

Search results for vLLM
vLLM community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including vLLM
vLLM integrated PyNvVideoCodec to offload video decoding from CPU to GPU's NVDEC. Result: 2x+ throughput at 8×H100, CPU bottleneck gone! This is huge for video captioning at scale (AV training, metadata), and ships with CUDA vLLM releases. 🔗
Show more
vLLM @vllm_project sessions at #PyTorchCon# NA 2026 span the serving stack, from attention and KV cache management to disaggregated serving and hardware portability across accelerators. Speakers will cover attention and KV cache systems, disaggregated serving, expert parallelism, and vLLM across TPU, Trainium, Arm, and IBM Spyre. Featured speakers include: @RedHat: Lucas Wilkinson, Matthew Bonanni, Zhanqiu Hu, @rickynds, and Alex Brooks @amazon: Sunita Nadampalli @IBM: Or Ozeri, Thomas Parnell, and Dave Grove @Huawei: Mengqing Cao @nvidia: Itay Alroy @Google: @Rob_Mulla and Qi Zhou @awscloud: Maen Suleiman @Meta: Richard Zou, Colin Taylor, and Angela Yi @BAAIBeijing: Yonghua Lin @fujitsulabs: Abhishek Jain and N Maajid Khan @IBMResearch: Antoni Viros i Martin and Avery Blanchard @MistralAI: Nicolò Lucchesi @tensormesh: @this_will_echo @googlecloud: Bill Jia Register by September 4 to save: Read the vLLM session guide:
Show more
vLLM Office Hours #56# recording is up: - What's new in vLLM 0.27 - Running Codex and Claude Code CLIs on open models - Speculators v0.6 and v0.7 update - GuideLLM v0.7.0: benchmarking reasoning, tool calling, and agentic traces Watch + see slides:
Show more
vLLM Conference is next week, and we have a packed schedule 🎊 📅. Here's the full list of events you should know: Mon: 🔷 4–6PM vLLM × Ray × Google Cloud Happy Hour: 🔷 6–9PM vLLM × Dynamo Meetup: Tue: 🔷 11–11:30AM vLLM Keynote from @simon_mo_ 🔷 12–5PM vLLM Track Day 1 🔷 6–9PM vLLM × AMD Happy Hour: Wed: 🔷 12–5PM vLLM Track Day 2 🔷 6–8:30PM vLLM x DigitalOcean × NVIDIA Happy Hour: No ticket needed for the happy hours and meetups, but space is limited. To join the full event, register here:
Show more
vLLM Office Hours today at 2pm ET: RL at 1T Scale, a prime-rl performance deep dive with @m_sirovatka (@PrimeIntellect). Training trillion-parameter MoE models like GLM-5.1 on agentic RL, plus what's new in vLLM 0.25 by @mgoin_. Get a recurring cal invite:
Show more
vLLM Office Hours is back Thursday 🗓️ Topic: Latest Trends in AI Agent Applications and @vllm_project 2:00 PM ET | 11:00 AM PT Join us live with @AIatAMD to dig into where agentic AI and vLLM intersect. Get a calendar invite:
Show more
vLLM tops the Artificial Analysis leaderboard 🎉 vLLM tops @ArtificialAnlys on DeepSeek V3.2 and ranks among the top deployments of MiniMax-M2.5 and Qwen 3.5 397B. The leading deployments of these models are now open source. How each result was built: 🔹 DeepSeek V3.2 — Aggressive op fusion across the attention path collapsed ~33 per-layer kernels down toward ~10. 🔹 MiniMax-M2.5 — Custom EAGLE3 draft trained against the target's own token distribution via TorchSpec, plus a custom QK-norm fusion for MiniMax's TP-aware attention. 🔹 Qwen 3.5 397B — Targeted fusions plus a QK-norm fix for Qwen's linear-attention path. Every optimization is in vLLM main or on its way upstream. Huge thank you to @inferact, @digitalocean, @nvidia, @RedHat_AI, and the vLLM community 🙏 Full breakdown 👇
Show more
🚀 vLLM's Humming backend can run Chord, @novita_labs' open-source W4A16 MoE kernels for Kimi K2.x. Kernel gains reach 1.33x on H200 TP8 vs tuned public Humming and 2.15x on B300 EP8 decode vs its untuned default. The indexed path works on compatible vLLM revisions; grouped integration is WIP. Great to see the kernels and benchmarks open-sourced! Details:
Show more
The vLLM blog is a very nice summary of the current status of speculative decoding methods (and has little info on AMD GPUs despite the title)
The @vllm_project & llm-d maintainers at @RedHat_AI are some of the hardest-working and leading experts on inference in the world 🚀 Everyone in the ML community can tell you how kind & helpful @robertshaw21 & @mgoin_ are. We’re grateful to occasionally collaborate with them. ❤️
Show more