🎉 Congrats to
@Alibaba_Qwen on Qwen3.8-2.4T-A95B, one of the largest open-weight models released to date. 2.4T params, 95B active, 512 experts.
Day-0 support in vLLM, verified on
@NVIDIA and
@AMD hardware. A ready-made 4-bit checkpoint per vendor, both out of the box:
Inferact/Qwen3.8-2.4T-A95B-NVFP4, 1.32 TiB, one NVIDIA 8xB300 node
Inferact/Qwen3.8-2.4T-A95B-MXFP4, 1.45 TiB, one AMD 8xMI355X node
No conversion, no calibration on your side. Just vllm serve.
Thanks to
@Alibaba_Qwen for the weights and the collaboration,
@NVIDIAAI and
@AIatAMD for the joint kernel engineering,
@inferact for the quantized checkpoints and vLLM integration,
@digitalocean and
@togethercompute for early testing, and the vLLM community. 🙌
🔗