Register and share your invite link to earn from video plays and referrals.

Search results for vLLM
vLLM community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including vLLM
vLLM Office Hours today at 2pm ET: RL at 1T Scale, a prime-rl performance deep dive with @m_sirovatka (@PrimeIntellect). Training trillion-parameter MoE models like GLM-5.1 on agentic RL, plus what's new in vLLM 0.25 by @mgoin_. Get a recurring cal invite:
Show more
vLLM Office Hours is back Thursday 🗓️ Topic: Latest Trends in AI Agent Applications and @vllm_project 2:00 PM ET | 11:00 AM PT Join us live with @AIatAMD to dig into where agentic AI and vLLM intersect. Get a calendar invite:
Show more
vLLM tops the Artificial Analysis leaderboard 🎉 vLLM tops @ArtificialAnlys on DeepSeek V3.2 and ranks among the top deployments of MiniMax-M2.5 and Qwen 3.5 397B. The leading deployments of these models are now open source. How each result was built: 🔹 DeepSeek V3.2 — Aggressive op fusion across the attention path collapsed ~33 per-layer kernels down toward ~10. 🔹 MiniMax-M2.5 — Custom EAGLE3 draft trained against the target's own token distribution via TorchSpec, plus a custom QK-norm fusion for MiniMax's TP-aware attention. 🔹 Qwen 3.5 397B — Targeted fusions plus a QK-norm fix for Qwen's linear-attention path. Every optimization is in vLLM main or on its way upstream. Huge thank you to @inferact, @digitalocean, @nvidia, @RedHat_AI, and the vLLM community 🙏 Full breakdown 👇
Show more
The @vllm_project & llm-d maintainers at @RedHat_AI are some of the hardest-working and leading experts on inference in the world 🚀 Everyone in the ML community can tell you how kind & helpful @robertshaw21 & @mgoin_ are. We’re grateful to occasionally collaborate with them. ❤️
Show more
If vLLM adopted a mascot….what would it be? Animal, object, character? Drop any ideas or suggestions below 👀
👀 vLLM community is working non-stop to get @deepseek_ai's new DSpark spec decode algorithm for vLLM! Faster inference for everyone!
🚀 vLLM-Omni v0.20.0 is out — aligned with upstream vLLM v0.20.0 (CUDA 13.0 · PyTorch 2.11 · Transformers 5.x). ⚡ Qwen3-Omni throughput +72% on H20, 32 conc (0.241 → 0.414 req/s) via talker / code2wav multi-replica scaling 🎙️ TTS faster & leaner: VoxCPM2 RTF 0.946 → 0.106 · Fish Speech Fast AR latency -53% · Qwen3-TTS / Voxtral-TTS Code2Wav saves ~3.2 GiB 🎨 Diffusion dynamic step-level batching: +7.8% throughput / -5.8% latency 🆕 New / improved: HunyuanImage-3.0, ERNIE T2I, AudioX, Wan2.2-S2V, LTX-2.3, FastGen Wan 2.1 📱 Wan2.2 on NPU production-ready: MindIE-SD, fused ops, VAE BF16, HSDP/USP — +50–60% perf 🧮 Quant expanded: Qwen Omni W4A16, OmniGen2 FP8, Z-Image FP8, HunyuanImage3 NPU, GLM-Image 🧩 Multi-backend updates across CUDA / ROCm / MUSA / NPU / XPU Check it out →
Show more
One of vLLM’s biggest advantages isn’t speed. It’s compatibility. Many applications can point existing OpenAI SDKs and API calls to vLLM with minimal code changes. That changes the conversation from just model quality to cost, control, and infrastructure flexibility.
Show more
weekend project: 2x3090/vllm cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 200k context. swival as my coding agent. As long as models keep getting more powerful via RL, distillation, and quantization, GPU depreciation will be much slower than expected. Even a 3090 will remain very useful
Show more
If you've wanted to learn vLLM but don't have a GPU sitting around, this is the path. Free Red Hat Developer Sandbox account. JupyterLab and models already deployed. You connect, run the notebooks, and learn by doing. No setup required. In ~1 hour: quantize a model, serve it via an OpenAI-compatible API, and benchmark it under real load. The same end-to-end workflow you'd run in production, without the infrastructure overhead. Start here:
Show more