๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Zhihu Frontier
@ZhihuFrontier
๐Ÿš€Bringing China's AI & tech trends, voices and perspectives to the global stage. โšก๏ธPowered by ็ŸฅไนŽ/ China's leading knowledge community.
๊ฐ€์ž… June 2025
180 ํŒ”๋กœ์ž‰ ์ค‘    11.6K ํŒฌ
๐Ÿ’ฅ GLM-5.3-Flash: Reported Opus 4.8-Level Scores at One-Tenth the Price Zhipu has confirmed that Ox Alpha โ€” the anonymous model that dominated OpenRouter and OpenCode last week โ€” is GLM-5.3-Flash, a 320B-parameter MoE with 18B active, released with open weights and an API at roughly one-tenth of GLM-5.3's price. Zhihu contributor ๅฐๅฐๅฐ† opens with a confession: he had hoped Ox Alpha was an external team fine-tuning the open GLM weights. He was wrong โ€” it was Zhipu's own model all along. His broader take: at this performance and this price, Flash becomes the new "kill line" for large models โ€” the bar below which rivals simply get priced out. 1๏ธโƒฃ The numbers behind the "kill line" Flash is somewhat larger than DeepSeek-V4 Flash, at 320B total and 18B active. On Zhipu's own benchmark comparisons, overall capability is roughly level with Claude Opus 4.8. More reference points he cites: ๐Ÿ”น 57 on the Artificial Analysis Intelligence Index โ€” around GPT-5.6 Terra's level, slightly below GLM-5.3. ๐Ÿ”น No.5 on Code Arena, one spot above GLM-5.3. ๐Ÿ”น API at 1/10 of GLM-5.3's price, with a limited-time half-off promo bringing it to 1/20. 2๏ธโƒฃ Same DeepSWE score, a fraction of the cost The author's sharpest comparison is on DeepSWE, where Flash scores 63% at a single-task cost of $0.24. DeepSeek-V4-Pro hits the same score at $1.67 per task โ€” roughly seven times the cost for equivalent results. This is the author's cost arithmetic on reported figures, not an independent measurement. 3๏ธโƒฃ The architecture that cuts the bill Part of the price drop is structural. Flash uses a hybrid of linear and sparse attention, sharply cutting attention compute, and a new IndexPool that compresses the indexer's cache from four copies to one โ€” reducing latency and memory at million-token context. Versus GLM-5.3, the company reports attention compute down to about 1/3 and KV cache down to about 1/4.4. 4๏ธโƒฃ The compute mystery, answered by domestic chips One reason few believed Ox Alpha was Zhipu's: the company was not thought to have enough spare compute for a massive free public test. The answer, per Zhipu: Flash is served from domestic Chinese chip clusters. The team built a custom inference engine on SGLang and worked around limited VRAM and bandwidth with quantization, layered deployment, and trading compute for communication โ€” lifting end-to-end serving performance about 3x on the same hardware, with per-token cost now close to mainstream NVIDIA GPUs. 5๏ธโƒฃ Chinese silicon just passed its biggest stress test The author's conclusion: Chinese chip clusters have shown they can carry large-scale inference for a frontier-level model โ€” including a free, record-breaking public trial. If that holds, he argues, the outlook for Chinese models just got a lot brighter. ๐Ÿ”— Full Reading: #GLM# #Zhipu# #GLM5Flash# #AIChips# #LLM# #AIInfra# #OpenWeights#
๋” ๋ณด๊ธฐ