註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Alexey Fateev
@superalesha
⚡I benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. ❤️ - 2xDGX Spark 🚀96GB VRAM | Local AI
加入 January 2026
325 正在關注    3.7K 粉絲
Im done with GLM-5.3 Flash 🤯 I went deep into the guts of GLM-5.3 Flash. Ripped out every expert, every memory head, every quantization scale, looked at all of it up close, then put the model back together so you can see the inside too. 🎉 GLM-5.3 Flash NVFP4 is now in Weight Atlas. 320B total, 18B active, 45 language layers and a 24 block Vision tower. And none of it is read from the config. I ran the deployed NVFP4 checkpoint itself on 4x RTX PRO 6000 and measured what it does. - 12,096 expert cells, 42 layers x 288 experts, each with exact REAP importance from 12.59M tokens per layer, route share and output contribution. Slice it by 14 domains, russian and vision included - one chart invalidated an assumption I had. Route count does not track expert importance, Spearman -0.222 against exact REAP. The ranking itself is stable, 0.994 split half - 2,176 KDA memory heads, each with its own half life, from 0.25 tokens to 2.7M - 19 billion NVFP4 block scales, plus the deployed FC2 QDQ error: median 21.05 dB and 21.7% of values turn into exact zeros - a causal check of REAP at inference, weights untouched. Dropping the top 2% experts changes 79% of sequences, dropping the bottom 2% changes 64% Every number comes from a capture I checksummed, 15 artifacts, and the limits are written next to each chart. Sound on I will record a video walk through every chart later. For now go poke around inside it:
顯示更多
0
11
94
5
轉發到社區