Qwen3.8-Flash-Next from
@Alibaba_Qwen has day-0 support in vLLM, verified on NVIDIA and AMD GPUs. đ
Ultra-sparse multimodal MoE: 125B params, 6B active, 262K native, 1M via YaRN. On top of those sits a separate 51B N-gram table you can offload.
Most of it will look familiar. The Gated DeltaNet layers reuse the KV path vLLM has had since Qwen3-Next: only a quarter of the layers hold a growing KV cache. Keep the 51B table in host RAM instead of HBM with VLLM_PLE_CPU_OFFLOAD=1. Qwen Sparse Attention is where the new engine work went. For now the model runs from vllm/vllm-openai:qwen38-flash-next.
Thanks to
@Alibaba_Qwen for the weights, and for opening them this early! đ
đ