๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Alexey Fateev
@superalesha
โšกI benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. โค๏ธ - 2xDGX Spark ๐Ÿš€96GB VRAM | Local AI
๊ฐ€์ž… January 2026
325 ํŒ”๋กœ์ž‰ ์ค‘    3.7K ํŒฌ
Im done with GLM-5.3 Flash ๐Ÿคฏ I went deep into the guts of GLM-5.3 Flash. Ripped out every expert, every memory head, every quantization scale, looked at all of it up close, then put the model back together so you can see the inside too. ๐ŸŽ‰ GLM-5.3 Flash NVFP4 is now in Weight Atlas. 320B total, 18B active, 45 language layers and a 24 block Vision tower. And none of it is read from the config. I ran the deployed NVFP4 checkpoint itself on 4x RTX PRO 6000 and measured what it does. - 12,096 expert cells, 42 layers x 288 experts, each with exact REAP importance from 12.59M tokens per layer, route share and output contribution. Slice it by 14 domains, russian and vision included - one chart invalidated an assumption I had. Route count does not track expert importance, Spearman -0.222 against exact REAP. The ranking itself is stable, 0.994 split half - 2,176 KDA memory heads, each with its own half life, from 0.25 tokens to 2.7M - 19 billion NVFP4 block scales, plus the deployed FC2 QDQ error: median 21.05 dB and 21.7% of values turn into exact zeros - a causal check of REAP at inference, weights untouched. Dropping the top 2% experts changes 79% of sequences, dropping the bottom 2% changes 64% Every number comes from a capture I checksummed, 15 artifacts, and the limits are written next to each chart. Sound on I will record a video walk through every chart later. For now go poke around inside it:
๋” ๋ณด๊ธฐ