Register and share your invite link to earn from video plays and referrals.

Search results for latepost
latepost community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including latepost
⚙ How GLM-5.3-Flash Served 70 Trillion Free Tokens on Chinese Chips GLM-5.3-Flash's anonymous "Ox Alpha" trial burned through roughly 70 trillion tokens in a week — and Zhipu says all of it ran on Chinese AI chip clusters. For many observers, that is a bigger story than the model itself. Zhihu contributor 恋猫 breaks down the systems engineering that made it work. His framing: the model architecture first reduces how much data needs to move, then the inference system compresses what remains. The reported scale is around 100,000 Chinese chips — from Huawei, Moore Threads and Hygon, according to @latepostnews — though @Zai_org itself has only said "tens of thousands." 1️⃣ The architecture cuts data movement first Versus GLM-5.3, Flash lowers attention compute by 3.01x and KV cache by 4.44x. Less data shuttling between memory and compute is the foundation everything else builds on. 2️⃣ Compute-for-bandwidth: trading FLOPs for HBM relief The bottleneck on these chips is HBM: the compute units still have headroom while memory bandwidth is nearly saturated. The fix, loosely speaking, turns "every step: read the full state, write the full state" into "read the state, recompute a little, write back periodically." A bit of extra matrix math buys a large reduction in HBM write traffic. 3️⃣ Communication-for-bandwidth: shard the cache across cards The cluster also uses high-speed inter-chip links and aggregated bandwidth to cut how much data each card must keep resident. The author's example: rank 0 holds the KV and indexer cache for some layers, rank 1 holds the rest. No card stores every layer's cache long-term; data is prefetched over the interconnect as each layer executes. The trade-off is real — more communication, far less resident cache per card. 4️⃣ EPD separation: three pools, scaled independently Finally, Encode, Prefill and Decode are split into three independently scalable resource pools, which matters most for multimodal traffic. Together, these pieces are what let the cluster run at high utilization. 5️⃣ The result: 3x end-to-end, cost "comparable to NVIDIA" Zhipu's own claim: end-to-end serving performance improved 3x on the same hardware, bringing per-token cost close to mainstream NVIDIA GPUs. The author's reading: individual Chinese cards may still be weaker, but model-system co-design lets the cluster as a whole reach international-mainstream throughput and cost. He adds a wry footnote: Zhipu's API used to be notorious among Chinese developers for 429 rate-limit errors. That it could absorb this launch's traffic at all — entirely on Chinese silicon — is, in his words, proof that optimization never ends. 🔗 Full Reading: #GLM# #Zhipu# #AIChips# #AIInfra# #LLM# #Inference# #OpenWeights#
Show more
🔌 GLM-5.3-Flash Served Its Viral Debut Entirely on Domestic Chinese Chips Zhipu's GLM-5.3-Flash — the 320B-A18B model revealed this week as the anonymous "Ox Alpha" — set usage records on OpenRouter and OpenCode during its undercover test. The company says all of that traffic was served by domestic Chinese chip clusters. Zhihu contributor 刘延 reconstructs how Zhipu lined up this infrastructure, and reads the official engineering details for hints about which chips are actually doing the work. The core judgment: the bigger story here is not the model itself, but that a frontier-level model handled global-scale, real-world inference on domestic silicon. 1️⃣ The timeline behind the launch The author pieces together a sequence from public reporting. 🔹 Zhipu was reported to have acquired an infrastructure company. 🔹 LatePost reported Zhipu had brought 50,000 domestic cards online; around the same time, its CodePlan subscription got cheaper with generous bonus quotas. 🔹 Ox Alpha went live anonymously and, in the author's words, blew up worldwide. 🔹 Zhipu then confirmed every request in that test ran on domestic chips. 🔹 The latest LatePost report puts the deployment at 100,000 domestic cards. Note the card counts come from media reports, not Zhipu itself. 2️⃣ The engineering: surviving 1M context on constrained hardware Zhipu's own statement is unusually specific about the constraints. The main bottleneck on these chips is memory capacity and bandwidth, and supporting a 1M-token context is the hardest part. The company's listed optimizations include trading compute for bandwidth and communication for memory, intra-node tensor parallelism for the linear attention and the LM head, ReplaySSM, W8A8 quantization, INT8/FP8/BF16 mixed cache quantization, and Layer Split. 3️⃣ Which chips? Reading the precision hints Here the author speculates, and it should be read as inference, not confirmation. 🔹 FP8 support suggests Moore Threads could be handling prefill, or possibly Hygon's DCU-3. 🔹 INT8 points toward Ascend 910B/C as the likely backbone. 🔹 No mention of FP4 suggests the newer Ascend 950 is probably not in the mix. 4️⃣ Why this matters If the reporting holds, this is the first time domestic Chinese chip clusters have carried a frontier model's global production traffic at this scale — including a free, record-breaking stress test from developers worldwide. The author treats it as a proof point: China's domestic chips are no longer just for training experiments or internal pilots, but can serve a top-tier model to the open internet. 🔗 Key links: Official announcement: Open weights (MIT): 🔗 Full Reading: #GLM# #Zhipu# #AIChips# #AIInfra# #Ascend# #Semiconductors# #OpenWeights#
Show more
Late post Santi v Corona 4 2/3 IP, 1 ER, 2SO, 4 Hits, FB sitting 90… & 3 for 3 hitting 2 singles, 1 double.