GLM-5.3-Flash, 320B params, one DGX Spark, 256K context, 33.8 tok/s peak.
Released 2 days ago
Serving recipe with MTP speculative decode already up, incl. the branch pick that makes it 2x faster:
GLM-5.3-Flash can now be run locally! ✨
Run 3-bit on 128GB RAM via Unsloth GGUF.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.
Guide:
GGUF: