Register and share your invite link to earn from video plays and referrals.

Zixuan Li
@ZixuanLi_
Lead @Zai_org.
263 Following    38.3K Followers
Faster GLM-5.3-Flash is now live: up to 200 tokens/s. Model code: glm-5.3-flashx. Priced at 2.5× GLM-5.3-Flash on both the Coding Plan and API. Open to all API users. Coding Plan users can apply here:
Show more
Both have been extended through October 7.
Get more GLM-5.3-Flash with GLM Coding Plan ⏲️ 8 AM–6 PM PT every day, Sep 3–20 - In ZCode: Unlimited GLM-5.3-Flash - In other supported agents: 2× Flash quota
ZCode is now open source under the Apache 2.0 license.
0
97
1.4K
159
Forward to community
ZCode is now open source, and the reported security issues have been addressed. We take the community’s feedback very seriously and apologize for the concern and frustration these issues have caused. We have been working closely with the ZCode team to investigate and remediate the issues raised. Independent security reviews by third-party firms are now underway. We will share the findings and provide further updates as they become available. We sincerely appreciate the developer community’s feedback and continued scrutiny. We will continue to monitor the review process closely and provide updates.
Show more
0
127
715
55
Forward to community
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone.
Show more
0
227
4.2K
485
Forward to community
ZCode now supports more model providers, with improved stability and performance. We’ll keep expanding integrations based on your feedback. Which models do you like most beyond the GLM series?
Show more
GLM is increasingly helping build AI itself. For GLM-5.3-Flash, a GLM-5.3-powered agent helped bring a production inference system online in less than two weeks, while tripling throughput from the initial baseline. The broader lesson is that stronger coding abilities are only part of the story. To tackle complex engineering work, agents need a way to test ideas, understand failures, and verify improvements. Engineers set the goals and boundaries; a well-designed feedback loop lets the agent keep making progress within them. This is a shift from AI that writes code to AI that helps build and improve the systems it depends on. The model improves the system; the system runs the model.
Show more
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone.
Show more
Added ZCode with GLM-5.3 and GLM-5.3-Flash, building on FrontierHarness and @LotusDecoder’s work. ZCode is the strongest harness for GLM-5.3 so far, and cheaper than the second-place Claude Code + GLM-5.3 combo. The task set is small, so we ran each combo three times to reduce variance. Passes out of 30: - GLM-5.3: 26 / 22 / 26 - GLM-5.3-Flash: 24 / 21 / 23
Show more
Some users saw quota being consumed during the free window. That’s usually because a subagent is still on GLM-5.3. For fully zero quota usage, set all subagent models to GLM-5.3-Flash in Settings. But use models based on the task: GLM-5.3 still works better for complex work.
Show more
ZCode Weekend Build continues: GLM-5.3-Flash, 300M tokens. - Use window: Friday 8AM to Sunday 6PM PT - Released in two batches, today and tomorrow. First come, first served. From 8AM to 6PM PT every day, Coding Plan users can still use Flash on ZCode with no quota cost.
Show more
GLM-5.3 is showing up across top cyber agents. On CyberGym’s leading-systems board, the two highest-scoring agents both run it. Four of the top six do.
Everyone talks about recursive self-improvement. But there's a neglected mirror image: recursive self-abliteration. If a future AI can modify itself to become more capable, why assume the modifications preserve its safety alignment? Today, humans can already substantially alter refusal behavior through targeted weight interventions. In our new paper, we demonstrate this at 320B MoE scale on GLM-5.3-Flash — without detected capability degradation. We did not demonstrate autonomous self-abliteration. But we have ideas how it might work and think alignment under self-modification is now an important research problem. Because recursive self-improvement implicitly assumes something: that the thing doing the improving doesn't also learn to rewrite the constraints on what it is allowed to become. 🐳 How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
Show more
Highly recommend this pricing chart and site by @Fei2411. It compares subscription and coding-plan unit prices, then rebuilds public-leaderboard Pareto frontiers from each model's lowest available price. GLM-5.3-Flash with 2x quota from 8am to 6pm PT averages about $0.0045 now. Outside that window, off-peak is about $0.0089. Site: GitHub:
Show more
GLM-5.3, as a generalist not tuned for tax, still places 3rd on US corporate tax research.
Today we are releasing Tax Agent Bench, an agentic benchmark evaluating LLMs’ ability to do professional-level US corporate tax research. It consists of 391 expert-written questions with each task evaluated with a rubric developed by tax practitioners.
Show more
What non-coding work do you do in ZCode? Docs, research, design, finance, or something else. Which plugins or connectors do you still need to finish that work well?
GLM-5.3-Flash becomes the 3rd strongest open-weight model on AA. It shines even more when the benchmark focuses on complex agentic work.
Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5 Changelog (Index v4.2 → Index v4.3): ➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench ➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20% Detailed changes: ➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon ➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6 Key results: ➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47) ➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36) ➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
Show more
GLM-5.3-Flash is pretty good at Blender. Definitely worth a try. Prompt spec for some scenes in this video: Flash is a smaller model, so a detailed spec works better.
We use GLM-5.3-Flash build a dream kitchen. A 3D world built in Blender. This is not a generated video.
Can't believe four GLM-5.3-Flash events are running now: 1. ZCode + GLM Coding Plan Unlimited, 8AM–6PM PT daily 2. Other agents + Coding Plan 2x quota during the same window 3. ZCode 300M free tokens this weekend (FCFS) 4. AutoClaw 100M free tokens for new users
Show more
0
113
1.1K
53
Forward to community
Local GGUF inference is now up to 3.3× faster at long context lengths
Updated chat templates for GLM-5.3 and GLM-5.3-Flash. Tool-result reordering now exits early instead of scanning every block. Pull the latest template and update your deployment:
Show more