註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Alok
@analogalok
Mechatronics Engineer AI belongs on your device. • Offline inference • No subscriptions. Teaching you to own your AI Intelligence Stack
加入 November 2013
223 正在關注    4.6K 粉絲
Ziphu achieved 3× performance gains and Nvidia level cost parity on domestic Chinese chips for serving GLM 3.5 Flash!. The wildest detail in the GLM-5.3-Flash release: Ziphu used a GLM 5.3 infrastructure agent to write custom GPU kernels, debug bottlenecks, and build an inference engine on top of SGLang, creating a self optimizing feedback loop that achieved 3× performance gains and Nvidia level cost parity on domestic Chinese chips.
顯示更多