Ziphu achieved 3× performance gains and Nvidia level cost parity on domestic Chinese chips for serving GLM 3.5 Flash!.
The wildest detail in the GLM-5.3-Flash release:
Ziphu used a GLM 5.3 infrastructure agent to write custom GPU kernels, debug bottlenecks, and build an inference engine on top of SGLang, creating a self optimizing feedback loop that achieved 3× performance gains and Nvidia level cost parity on domestic Chinese chips.