注册并分享邀请链接,可获得视频播放与邀请奖励。

Zhihu Frontier
@ZhihuFrontier
🚀Bringing China's AI & tech trends, voices and perspectives to the global stage. ⚡️Powered by 知乎/ China's leading knowledge community.
加入 June 2025
180 正在关注    11.6K 粉丝
💥 GLM-5.3-Flash: Reported Opus 4.8-Level Scores at One-Tenth the Price Zhipu has confirmed that Ox Alpha — the anonymous model that dominated OpenRouter and OpenCode last week — is GLM-5.3-Flash, a 320B-parameter MoE with 18B active, released with open weights and an API at roughly one-tenth of GLM-5.3's price. Zhihu contributor 小小将 opens with a confession: he had hoped Ox Alpha was an external team fine-tuning the open GLM weights. He was wrong — it was Zhipu's own model all along. His broader take: at this performance and this price, Flash becomes the new "kill line" for large models — the bar below which rivals simply get priced out. 1️⃣ The numbers behind the "kill line" Flash is somewhat larger than DeepSeek-V4 Flash, at 320B total and 18B active. On Zhipu's own benchmark comparisons, overall capability is roughly level with Claude Opus 4.8. More reference points he cites: 🔹 57 on the Artificial Analysis Intelligence Index — around GPT-5.6 Terra's level, slightly below GLM-5.3. 🔹 No.5 on Code Arena, one spot above GLM-5.3. 🔹 API at 1/10 of GLM-5.3's price, with a limited-time half-off promo bringing it to 1/20. 2️⃣ Same DeepSWE score, a fraction of the cost The author's sharpest comparison is on DeepSWE, where Flash scores 63% at a single-task cost of $0.24. DeepSeek-V4-Pro hits the same score at $1.67 per task — roughly seven times the cost for equivalent results. This is the author's cost arithmetic on reported figures, not an independent measurement. 3️⃣ The architecture that cuts the bill Part of the price drop is structural. Flash uses a hybrid of linear and sparse attention, sharply cutting attention compute, and a new IndexPool that compresses the indexer's cache from four copies to one — reducing latency and memory at million-token context. Versus GLM-5.3, the company reports attention compute down to about 1/3 and KV cache down to about 1/4.4. 4️⃣ The compute mystery, answered by domestic chips One reason few believed Ox Alpha was Zhipu's: the company was not thought to have enough spare compute for a massive free public test. The answer, per Zhipu: Flash is served from domestic Chinese chip clusters. The team built a custom inference engine on SGLang and worked around limited VRAM and bandwidth with quantization, layered deployment, and trading compute for communication — lifting end-to-end serving performance about 3x on the same hardware, with per-token cost now close to mainstream NVIDIA GPUs. 5️⃣ Chinese silicon just passed its biggest stress test The author's conclusion: Chinese chip clusters have shown they can carry large-scale inference for a frontier-level model — including a free, record-breaking public trial. If that holds, he argues, the outlook for Chinese models just got a lot brighter. 🔗 Full Reading: #GLM# #Zhipu# #GLM5Flash# #AIChips# #LLM# #AIInfra# #OpenWeights#
显示更多