💥 GLM-5.3-Flash: Reported Opus 4.8-Level Scores at One-Tenth the Price
Zhipu has confirmed that Ox Alpha — the anonymous model that dominated OpenRouter and OpenCode last week — is GLM-5.3-Flash, a 320B-parameter MoE with 18B active, released with open weights and an API at roughly one-tenth of GLM-5.3's price.
Zhihu contributor 小小将 opens with a confession: he had hoped Ox Alpha was an external team fine-tuning the open GLM weights. He was wrong — it was Zhipu's own model all along.
His broader take: at this performance and this price, Flash becomes the new "kill line" for large models — the bar below which rivals simply get priced out.
1️⃣ The numbers behind the "kill line"
Flash is somewhat larger than DeepSeek-V4 Flash, at 320B total and 18B active. On Zhipu's own benchmark comparisons, overall capability is roughly level with Claude Opus 4.8.
More reference points he cites:
🔹 57 on the Artificial Analysis Intelligence Index — around GPT-5.6 Terra's level, slightly below GLM-5.3.
🔹 No.5 on Code Arena, one spot above GLM-5.3.
🔹 API at 1/10 of GLM-5.3's price, with a limited-time half-off promo bringing it to 1/20.
2️⃣ Same DeepSWE score, a fraction of the cost
The author's sharpest comparison is on DeepSWE, where Flash scores 63% at a single-task cost of $0.24.
DeepSeek-V4-Pro hits the same score at $1.67 per task — roughly seven times the cost for equivalent results. This is the author's cost arithmetic on reported figures, not an independent measurement.
3️⃣ The architecture that cuts the bill
Part of the price drop is structural. Flash uses a hybrid of linear and sparse attention, sharply cutting attention compute, and a new IndexPool that compresses the indexer's cache from four copies to one — reducing latency and memory at million-token context.
Versus GLM-5.3, the company reports attention compute down to about 1/3 and KV cache down to about 1/4.4.
4️⃣ The compute mystery, answered by domestic chips
One reason few believed Ox Alpha was Zhipu's: the company was not thought to have enough spare compute for a massive free public test.
The answer, per Zhipu: Flash is served from domestic Chinese chip clusters. The team built a custom inference engine on SGLang and worked around limited VRAM and bandwidth with quantization, layered deployment, and trading compute for communication — lifting end-to-end serving performance about 3x on the same hardware, with per-token cost now close to mainstream NVIDIA GPUs.
5️⃣ Chinese silicon just passed its biggest stress test
The author's conclusion: Chinese chip clusters have shown they can carry large-scale inference for a frontier-level model — including a free, record-breaking public trial.
If that holds, he argues, the outlook for Chinese models just got a lot brighter.
🔗 Full Reading:
#
GLM# #
Zhipu# #
GLM5Flash# #
AIChips# #
LLM# #
AIInfra# #
OpenWeights#