Day zero, and we're already running it.
GLM-5.3-Flash: 320B params, 18B active. Frontier intelligence at flash-tier cost.
Efficient architecture needs compute that scales.
That's where comes in.
More capable and affordable models and accessible compute is where the AI industry is headed. And we're helping to make it happen.
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog:
Available now across all official platforms:
Weights:
API:
Coding Plan:
ZCode:
Chat:
AutoClaw: