註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    46.7K 粉絲
5x lower token costs on DeepSeek V4 in one month! Highlighting the vLLM community at work: day-zero recipes, then relentless optimization across kernels, scheduling, and serving. Every PR counts. 🚀
顯示更多
💡 Continuous software innovation is the force multiplier behind AI infrastructure — compounding inference performance, lowering cost per token, and increasing long-term value with every optimization. Open source accelerates this advantage. Leading AI frameworks like @PyTorch and inference engines such as @sgl_project and @vllm_project are built natively on NVIDIA CUDA, enabling research breakthroughs and software optimizations to unlock great performance on NVIDIA GPUs from day zero.
顯示更多