註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    50.3K 粉絲
Qwen3.8-Flash-Next from @Alibaba_Qwen has day-0 support in vLLM, verified on NVIDIA and AMD GPUs. 🎉 Ultra-sparse multimodal MoE: 125B params, 6B active, 262K native, 1M via YaRN. On top of those sits a separate 51B N-gram table you can offload. Most of it will look familiar. The Gated DeltaNet layers reuse the KV path vLLM has had since Qwen3-Next: only a quarter of the layers hold a growing KV cache. Keep the 51B table in host RAM instead of HBM with VLLM_PLE_CPU_OFFLOAD=1. Qwen Sparse Attention is where the new engine work went. For now the model runs from vllm/vllm-openai:qwen38-flash-next. Thanks to @Alibaba_Qwen for the weights, and for opening them this early! 🙌 🔗
顯示更多
0
36
374
53
轉發到社區