🔌 GLM-5.3-Flash Served Its Viral Debut Entirely on Domestic Chinese Chips
Zhipu's GLM-5.3-Flash — the 320B-A18B model revealed this week as the anonymous "Ox Alpha" — set usage records on OpenRouter and OpenCode during its undercover test. The company says all of that traffic was served by domestic Chinese chip clusters.
Zhihu contributor 刘延 reconstructs how Zhipu lined up this infrastructure, and reads the official engineering details for hints about which chips are actually doing the work.
The core judgment: the bigger story here is not the model itself, but that a frontier-level model handled global-scale, real-world inference on domestic silicon.
1️⃣ The timeline behind the launch
The author pieces together a sequence from public reporting.
🔹 Zhipu was reported to have acquired an infrastructure company.
🔹 LatePost reported Zhipu had brought 50,000 domestic cards online; around the same time, its CodePlan subscription got cheaper with generous bonus quotas.
🔹 Ox Alpha went live anonymously and, in the author's words, blew up worldwide.
🔹 Zhipu then confirmed every request in that test ran on domestic chips.
🔹 The latest LatePost report puts the deployment at 100,000 domestic cards.
Note the card counts come from media reports, not Zhipu itself.
2️⃣ The engineering: surviving 1M context on constrained hardware
Zhipu's own statement is unusually specific about the constraints. The main bottleneck on these chips is memory capacity and bandwidth, and supporting a 1M-token context is the hardest part.
The company's listed optimizations include trading compute for bandwidth and communication for memory, intra-node tensor parallelism for the linear attention and the LM head, ReplaySSM, W8A8 quantization, INT8/FP8/BF16 mixed cache quantization, and Layer Split.
3️⃣ Which chips? Reading the precision hints
Here the author speculates, and it should be read as inference, not confirmation.
🔹 FP8 support suggests Moore Threads could be handling prefill, or possibly Hygon's DCU-3.
🔹 INT8 points toward Ascend 910B/C as the likely backbone.
🔹 No mention of FP4 suggests the newer Ascend 950 is probably not in the mix.
4️⃣ Why this matters
If the reporting holds, this is the first time domestic Chinese chip clusters have carried a frontier model's global production traffic at this scale — including a free, record-breaking stress test from developers worldwide.
The author treats it as a proof point: China's domestic chips are no longer just for training experiments or internal pilots, but can serve a top-tier model to the open internet.
🔗 Key links:
Official announcement:
Open weights (MIT):
🔗 Full Reading:
#
GLM# #
Zhipu# #
AIChips# #
AIInfra# #
Ascend# #
Semiconductors# #
OpenWeights#