GLM-5.3-Flash (320B-A18B) from
@Zai_org drops today with day-0 support in SGLang. You may know it as ox-alpha from the past few days!
đ It's the first native multimodal model in the GLM-5 series, able to review its own output visually and fix what's wrong.
đ It outperforms GLM-5.2 at 1/10 the cost. Hybrid sparse plus linear attention enables stable 1M long-context performance at extreme cost efficiency.
đģ It goes beyond coding into professional work: slides, documents, spreadsheets, and finance research, all handled end to end.
Run GLM-5.3-Flash with SGLang, and welcome to a new era of efficient, production-ready open intelligence!