Register and share your invite link to earn from video plays and referrals.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
Joined March 2024
36 Following    45.4K Followers
5x lower token costs on DeepSeek V4 in one month! Highlighting the vLLM community at work: day-zero recipes, then relentless optimization across kernels, scheduling, and serving. Every PR counts. 🚀
Show more
💡 Continuous software innovation is the force multiplier behind AI infrastructure — compounding inference performance, lowering cost per token, and increasing long-term value with every optimization. Open source accelerates this advantage. Leading AI frameworks like @PyTorch and inference engines such as @sgl_project and @vllm_project are built natively on NVIDIA CUDA, enabling research breakthroughs and software optimizations to unlock great performance on NVIDIA GPUs from day zero.
Show more