Register and share your invite link to earn from video plays and referrals.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
Joined March 2024
36 Following    50.2K Followers
🎉 Congrats to @Alibaba_Qwen on Qwen3.8-2.4T-A95B, one of the largest open-weight models released to date. 2.4T params, 95B active, 512 experts. Day-0 support in vLLM, verified on @NVIDIA and @AMD hardware. A ready-made 4-bit checkpoint per vendor, both out of the box: Inferact/Qwen3.8-2.4T-A95B-NVFP4, 1.32 TiB, one NVIDIA 8xB300 node Inferact/Qwen3.8-2.4T-A95B-MXFP4, 1.45 TiB, one AMD 8xMI355X node No conversion, no calibration on your side. Just vllm serve. Thanks to @Alibaba_Qwen for the weights and the collaboration, @NVIDIAAI and @AIatAMD for the joint kernel engineering, @inferact for the quantized checkpoints and vLLM integration, @digitalocean and @togethercompute for early testing, and the vLLM community. 🙌 🔗
Show more