Register and share your invite link to earn from video plays and referrals.

Search results for Zhipu
Zhipu community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Zhipu
Zhipu is testing the same assumption DeepSeek tested a year and a half ago: that American frontier labs have an uncatchable lead. Zhipu +15%, Minimax (another Chinese AI lab) +24% overnight in HK While US AI names are falling today: Google -6.5%, Amazon -4.3%, Meta -2.6% @satyanadella summed up the shift:
Show more
China's Zhipu AI plans to apply for Shanghai's Sci-Tech Board listing
🚨 BREAKING: Anthropic publicly accused Zhipu's GLM-5.2 of distilling Claude and OpenAI models,
The founder of Chinese AI lab Zhipu argued that frontier AI should remain broadly accessible rather than controlled by select individuals, weighing in on a growing debate about the risks posed by ever more powerful models.
Show more
JUST IN: 🇨🇳 China’s Zhipu AI reportedly matches Anthropic’s Claude Mythos in security bug detection performance​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​.
Show more
0
257
4.6K
366
Forward to community
New model drop in Sider: GLM-5 is here. Zhipu AI's latest flagship brings 744B parameters, a 200K context window, and frontier-level coding & reasoning, now live in your sidebar. One more reason you'll never need to leave your browser.
Show more
⚔ GLM-5.3-Flash Aims Straight at the Post-Hike DeepSeek Zhipu's GLM-5.3-Flash — revealed this week as the anonymous "Ox Alpha" — has been open-weighted and priced at roughly one-tenth of GLM-5.3. Much of the early discussion compares it to DeepSeek's V4 Flash, which recently raised prices. Zhihu contributor 起步十档, who ran Ox Alpha inside real workflows before the reveal, gives a practitioner's verdict in one line: it is built to kill the post-hike DeepSeek. His case rests on three legs — performance, token efficiency, and an architecture change that makes the price possible. 1️⃣ It clears the bar for long-horizon work Official scores put Flash between Grok 4.6 and GLM-5.3, and clearly ahead of DeepSeek V4 Flash. In the author's own testing, its frontend ability roughly matches an early gray-test build of DeepSeek V4 Pro, while its backend is noticeably weaker than GLM-5.3 — but still usable on long-horizon tasks as long as the connection holds. His rule of thumb: any model past the Claude Opus 4.6 line is workflow-ready for long tasks. Beyond that, differences come down to reasoning style and accuracy, not viability. One caveat he flags: during the anonymous test the deployment was unstable, and some believe it served a mid-training checkpoint rather than the final model. 2️⃣ The real weapon: token efficiency Comparing peak API prices against the post-hike DeepSeek V4 Flash, the author notes cached input is actually 2x more expensive, while regular input and output sit at roughly 26% of DeepSeek's price. Since cached input is often the bulk of the bill, he wants real-world tests before calling a winner on price alone. But his own usage points the same direction. In one to two hours of real work — reading and writing files, running tests — Ox Alpha burned barely over 100K tokens. He estimates DeepSeek would need 250-300K for the same workload. His prediction: same tasks, run on both APIs, will come out cheaper on Flash — with clearly better performance. 3️⃣ His unexpected advice: skip the Coding Plan The plan only triples your quota for Flash. Using Zhipu's own best-case math — maximum usage, off-peak hours, the official 0.8x API-equivalent rate — the plan works out to about 53% of pay-as-you-go API cost. Since that scenario is already extreme, he concludes the API is the better deal for almost everyone. 4️⃣ The architecture change behind the price From GLM-5 through 5.3, Zhipu used DSA — essentially an optimized full attention — which costs more than DeepSeek's CSA/HCA, Kimi's KDA, or Qwen's linear-global hybrid. That is why GLM used to be pricier than larger DeepSeek models. Flash is the first GLM to switch to linear attention plus an HCA-like compressed attention, trained with the HCA recipe as well. That brings it in line with mainstream domestic practice — and the price fell accordingly. The author expects a future GLM-5.5 can scale up without costing much more than 5.3. On top of that, Zhipu added native multimodality and leaned on domestic compute, which he reads as the reason for the generous free quotas during the anonymous test. 5️⃣ Zhipu is still the team to beat The author's closing line is unambiguous: Zhipu remains, in his words, the number-one Chinese model company. That is his judgment, not a benchmark result — but the cost argument underneath it is now easy to check yourself. 🔗 Key links: Official announcement: Open weights (MIT): 🔗 Full Reading: #GLM# #Zhipu# #DeepSeek# #LLM# #AIInference# #TokenEfficiency# #OpenWeights#
Show more
🔌 GLM-5.3-Flash Served Its Viral Debut Entirely on Domestic Chinese Chips Zhipu's GLM-5.3-Flash — the 320B-A18B model revealed this week as the anonymous "Ox Alpha" — set usage records on OpenRouter and OpenCode during its undercover test. The company says all of that traffic was served by domestic Chinese chip clusters. Zhihu contributor 刘延 reconstructs how Zhipu lined up this infrastructure, and reads the official engineering details for hints about which chips are actually doing the work. The core judgment: the bigger story here is not the model itself, but that a frontier-level model handled global-scale, real-world inference on domestic silicon. 1️⃣ The timeline behind the launch The author pieces together a sequence from public reporting. 🔹 Zhipu was reported to have acquired an infrastructure company. 🔹 LatePost reported Zhipu had brought 50,000 domestic cards online; around the same time, its CodePlan subscription got cheaper with generous bonus quotas. 🔹 Ox Alpha went live anonymously and, in the author's words, blew up worldwide. 🔹 Zhipu then confirmed every request in that test ran on domestic chips. 🔹 The latest LatePost report puts the deployment at 100,000 domestic cards. Note the card counts come from media reports, not Zhipu itself. 2️⃣ The engineering: surviving 1M context on constrained hardware Zhipu's own statement is unusually specific about the constraints. The main bottleneck on these chips is memory capacity and bandwidth, and supporting a 1M-token context is the hardest part. The company's listed optimizations include trading compute for bandwidth and communication for memory, intra-node tensor parallelism for the linear attention and the LM head, ReplaySSM, W8A8 quantization, INT8/FP8/BF16 mixed cache quantization, and Layer Split. 3️⃣ Which chips? Reading the precision hints Here the author speculates, and it should be read as inference, not confirmation. 🔹 FP8 support suggests Moore Threads could be handling prefill, or possibly Hygon's DCU-3. 🔹 INT8 points toward Ascend 910B/C as the likely backbone. 🔹 No mention of FP4 suggests the newer Ascend 950 is probably not in the mix. 4️⃣ Why this matters If the reporting holds, this is the first time domestic Chinese chip clusters have carried a frontier model's global production traffic at this scale — including a free, record-breaking stress test from developers worldwide. The author treats it as a proof point: China's domestic chips are no longer just for training experiments or internal pilots, but can serve a top-tier model to the open internet. 🔗 Key links: Official announcement: Open weights (MIT): 🔗 Full Reading: #GLM# #Zhipu# #AIChips# #AIInfra# #Ascend# #Semiconductors# #OpenWeights#
Show more
⚡ GLM-5.3-Flash: Near-Flagship Logic at One-Tenth the Price Zhipu has confirmed that Ox Alpha — the anonymous model that just topped usage charts on OpenRouter and OpenCode — is GLM-5.3-Flash, a 320B-parameter MoE with 18B active, released with open weights. The company prices its API at roughly one-tenth of GLM-5.3, and reports an Artificial Analysis Intelligence Index score of 57, on par with Claude Opus 4.8. Zhihu contributor toyama nao, known for a long-running monthly logic benchmark built on self-designed problem sets, argues the launch fills a gap the market has had since DeepSeek raised prices: a model good enough to use daily and cheap enough to ignore. The core judgment: GLM-5.3-Flash is not a capability breakthrough. It is a cost breakthrough — same lineage as GLM-5.3, nearly identical results on many tasks, but a new path on inference efficiency. 1️⃣ Why the market needed this model After DeepSeek's price hike, its tier lost any clear price-performance leader. GLM-5.3 then improved quality without raising prices and quietly took that spot. Given compute costs in China, shrinking the model is the pragmatic route to genuinely low prices — and smaller models have already proven they can carry real workloads. What the market lacked was a model cheap enough that cost stops being a decision. Flash is that option. 2️⃣ Where Flash matches GLM-5.3 — and where the floor drops In the author's monthly logic evaluation, Flash matches the standard GLM-5.3 on most low- and mid-difficulty tasks, including coding. On hard tasks it can still reach the same ceiling — just not reliably. In practice that means more retries to get the best answer. Retries are cheap at Flash's pricing, but output speed hasn't improved, so the experience still degrades. 3️⃣ Hallucination: wider variance in both directions Flash's hallucination behavior shows a wider spread than the standard model. 🔹 At its best, it hallucinates less than GLM-5.3, catching very fine details buried in the context. 🔹 At its worst, it is worse — misreading long prompts and making basic mistakes. Longer inputs trigger the bad case more often. 4️⃣ The real story: token efficiency without longer reasoning Unlike many small models that stretch their reasoning chains to buy intelligence, Flash uses fewer tokens than GLM-5.3 on nearly every task — as little as 40% of the standard model's consumption in the best cases. On problems where most models brute-force the search space, Flash often narrows it down almost by intuition. The author notes GPT-5.6 still holds the token-efficiency crown on some difficulty levels. One catch: on tasks that genuinely require exhaustive search, Flash's consumption matches the standard model — and it can hit the API's 128K max output cap, leaving answers truncated. 5️⃣ Why this matters Four years into the LLM era, most people have still never used one for daily work. Products built on top of these models urgently need a model cheap enough to flip their ROI positive. As flagships cross the "good enough" threshold and their capability starts to overflow, that surplus intelligence should be inherited by a more civilian model. The market needed a low-price model — someone had to ship it. 🔗 Full Reading: 📊Author's monthly logic benchmark: #GLM# #Zhipu# #GLM5Flash# #LLM# #AIInference# #OpenWeights# #AIAgents#
Show more
Ox Alpha just got exposed.💀 Ox Alpha : Technical fingerprints are reportedly pointing toward Zhipu’s unreleased GLM. What we know so far: - Hit 80% on an early DeepSWE test - outperforming of Fable 5 and GPT-5.6 Sol - Reportedly 100T tokens/day capacity - Still offering 1M context & video input - Already being heavily used by coding agents Wildest part? It might not even be a new model just a stealth public test of Zhipu’s next flagship. Could this be our first real look at the next GLM?
Show more