With our new efficiency methods, you will be able to run this on a single DGX Spark or AMD Strix Halo at 7 token/s decode and >250 tok/s prefill. Stay tuned!
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense.
- Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model
- A major leap in cybersecurity, setting a new standard among open models
Tech Blog: