We made GLM-5.3-Flash run 3.3x faster locally!
Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction.
Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp.
Guide:
GGUF:
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog:
Available now across all official platforms:
Weights:
API:
Coding Plan:
ZCode:
Chat:
AutoClaw: