突发:阿里把 Qwen4 的架构提前开源了!
名字叫 Qwen3.8-Flash-Next, 今天刚放权重。
GDN + QSA 混合注意力、Gated Residual、N-gram Embedding、Muon 优化器——官方自己说这套就是 Qwen4 的前身。
参数账很怪 : 125B 主体 + 51B N-gram 表 , 每个 token 只激活 6B。
官方口径 : 训练成本只有 Qwen3.7-Plus 的 1/9 , 成绩反而全面更高。
SWE-bench Pro 62.5——12 天前刚开源的 27B 是 61.7, 被自家这个 6B 激活量的模型压过去了。
262K 原生上下文,可扩到 1M, 多模态。
API 定价输入 $0.16/1M、输出 $0.47/1M。
我更关心能不能本地跑。
目前的进展:
llama.cpp 支持已经在 PR(#
27739#) , mmproj、MTP 都有,还专门给那张 51B 的表做了 RAM offload——有人拆过 , 表有 3.2 亿行, 但每个 token 只查 16 行, 大头放内存就行,显卡只管 6B 激活。
已经有人算完账 , 判断 4 张 3090 就能跑。
等 GGUF 出来, 我会在 4090 48G 上把它和 27B 对着测一轮
官方发布👇
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency.
What's new: 🥳
- Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4.
- Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks.
- Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI).
- 262K native context, extensible to 1M with YaRN.
We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀
We can't wait to see what you build with Qwen3.8-Flash!👀👇
- Blog:
- Technical Report:
- Hugging Face:
- ModelScope:
もっと見る