⚔ GLM-5.3-Flash Aims Straight at the Post-Hike DeepSeek
Zhipu's GLM-5.3-Flash — revealed this week as the anonymous "Ox Alpha" — has been open-weighted and priced at roughly one-tenth of GLM-5.3. Much of the early discussion compares it to DeepSeek's V4 Flash, which recently raised prices.
Zhihu contributor 起步十档, who ran Ox Alpha inside real workflows before the reveal, gives a practitioner's verdict in one line: it is built to kill the post-hike DeepSeek.
His case rests on three legs — performance, token efficiency, and an architecture change that makes the price possible.
1️⃣ It clears the bar for long-horizon work
Official scores put Flash between Grok 4.6 and GLM-5.3, and clearly ahead of DeepSeek V4 Flash. In the author's own testing, its frontend ability roughly matches an early gray-test build of DeepSeek V4 Pro, while its backend is noticeably weaker than GLM-5.3 — but still usable on long-horizon tasks as long as the connection holds.
His rule of thumb: any model past the Claude Opus 4.6 line is workflow-ready for long tasks. Beyond that, differences come down to reasoning style and accuracy, not viability.
One caveat he flags: during the anonymous test the deployment was unstable, and some believe it served a mid-training checkpoint rather than the final model.
2️⃣ The real weapon: token efficiency
Comparing peak API prices against the post-hike DeepSeek V4 Flash, the author notes cached input is actually 2x more expensive, while regular input and output sit at roughly 26% of DeepSeek's price. Since cached input is often the bulk of the bill, he wants real-world tests before calling a winner on price alone.
But his own usage points the same direction. In one to two hours of real work — reading and writing files, running tests — Ox Alpha burned barely over 100K tokens. He estimates DeepSeek would need 250-300K for the same workload.
His prediction: same tasks, run on both APIs, will come out cheaper on Flash — with clearly better performance.
3️⃣ His unexpected advice: skip the Coding Plan
The plan only triples your quota for Flash. Using Zhipu's own best-case math — maximum usage, off-peak hours, the official 0.8x API-equivalent rate — the plan works out to about 53% of pay-as-you-go API cost.
Since that scenario is already extreme, he concludes the API is the better deal for almost everyone.
4️⃣ The architecture change behind the price
From GLM-5 through 5.3, Zhipu used DSA — essentially an optimized full attention — which costs more than DeepSeek's CSA/HCA, Kimi's KDA, or Qwen's linear-global hybrid. That is why GLM used to be pricier than larger DeepSeek models.
Flash is the first GLM to switch to linear attention plus an HCA-like compressed attention, trained with the HCA recipe as well. That brings it in line with mainstream domestic practice — and the price fell accordingly. The author expects a future GLM-5.5 can scale up without costing much more than 5.3.
On top of that, Zhipu added native multimodality and leaned on domestic compute, which he reads as the reason for the generous free quotas during the anonymous test.
5️⃣ Zhipu is still the team to beat
The author's closing line is unambiguous: Zhipu remains, in his words, the number-one Chinese model company. That is his judgment, not a benchmark result — but the cost argument underneath it is now easy to check yourself.
🔗 Key links:
Official announcement:
Open weights (MIT):
🔗 Full Reading:
#
GLM# #
Zhipu# #
DeepSeek# #
LLM# #
AIInference# #
TokenEfficiency# #
OpenWeights#