🚨 GLM-5.3-Flash just landed on LobeHub, and the architecture is kind of insane.
It has roughly the same total parameter count as GLM-4.5:
355B → 320B total params
But under the hood:
→ Active params: 32B → 18B
→ Layers: 92 → 45
→ 30T-token multimodal pretraining corpus
→ 1M context window
→ Native multimodality + tool calling
Basically,
@Zai_org cut the compute almost in half and somehow made the model stronger.
That’s why GLM-5.3-Flash can deliver frontier-level capability at a dramatically lower cost.
This is what the next model war looks like: not just bigger models, but much more intelligence per dollar.
Now available on LobeHub. ⚡