注册并分享邀请链接,可获得视频播放与邀请奖励。

Alok
@analogalok
Mechatronics Engineer AI belongs on your device. • Offline inference • No subscriptions. Teaching you to own your AI Intelligence Stack
加入 November 2013
223 正在关注    4.6K 粉丝
Ziphu achieved 3× performance gains and Nvidia level cost parity on domestic Chinese chips for serving GLM 3.5 Flash!. The wildest detail in the GLM-5.3-Flash release: Ziphu used a GLM 5.3 infrastructure agent to write custom GPU kernels, debug bottlenecks, and build an inference engine on top of SGLang, creating a self optimizing feedback loop that achieved 3× performance gains and Nvidia level cost parity on domestic Chinese chips.
显示更多