Im done with GLM-5.3 Flash 🤯
I went deep into the guts of GLM-5.3 Flash. Ripped out every expert, every memory head, every quantization scale, looked at all of it up close, then put the model back together so you can see the inside too.
🎉 GLM-5.3 Flash NVFP4 is now in Weight Atlas. 320B total, 18B active, 45 language layers and a 24 block Vision tower. And none of it is read from the config. I ran the deployed NVFP4 checkpoint itself on 4x RTX PRO 6000 and measured what it does.
- 12,096 expert cells, 42 layers x 288 experts, each with exact REAP importance from 12.59M tokens per layer, route share and output contribution. Slice it by 14 domains, russian and vision included
- one chart invalidated an assumption I had. Route count does not track expert importance, Spearman -0.222 against exact REAP. The ranking itself is stable, 0.994 split half
- 2,176 KDA memory heads, each with its own half life, from 0.25 tokens to 2.7M
- 19 billion NVFP4 block scales, plus the deployed FC2 QDQ error: median 21.05 dB and 21.7% of values turn into exact zeros
- a causal check of REAP at inference, weights untouched. Dropping the top 2% experts changes 79% of sequences, dropping the bottom 2% changes 64%
Every number comes from a capture I checksummed, 15 artifacts, and the limits are written next to each chart.
Sound on
I will record a video walk through every chart later. For now go poke around inside it:
顯示更多