Yesterday Xiaomi. Today Zhipu - sharing training process (and potential failure!) - the scarce skill is no longer pretrain. It’s scaling agentic RL + standing up serving on constrained hardware.
Xiaomi is livestreaming MiMo-V2.6 mid-run: ~2B tokens/step, 1568 prompts × 16 rollouts, fully async, mixed harnesses, extra compute on graders/credit assignment.
@Zai_org ai used GLM-5.3 as an infra agent to take GLM-5.3-Flash from first successful run to production in <2 weeks and 3× e2e throughput via traces, microbenchmarks, and local correctness tests- including on Chinese chips.
They’re publishing the method because that’s the moat they can prove. Weights and dashboards recruit talent and adoption. Secret pretrain recipes don’t, once everyone is already at the frontier on post-training. So ballsy!!
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash.
The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline.
The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone.
Show more