๐คฏ A 600B parameter model is going open-weight Oct. 15โฆ
โฆbut only 27B parameters are active per token. ๐๐
Here is Step 5 Preview (this model is exciting) and here is why โฆ
Stats ๐
๐ง 600B total parameters
โก 27B active/token
๐ 1M context
๐๏ธ vision + video
๐ค built for long-running agents
๐ open weights Oct. 15
And that 27B-active number makes this really interesting for Local AI right?
At roughly 4-bit, 600B parameters would theoretically be ~300GB of weights.
Real-world quantization + overhead means Iโd expect something more like ~320โ350GB before accounting for KV cache and other runtime memory.
So weโre potentially looking at
๐พ ~384GB-class memory โ Q4 territory
๐พ ~256GB-class memory โ aggressive Q3 territory
And because this is MoE, only 27B parameters participate in each tokenโs compute.
โ ๏ธ That does NOT mean this is a 27B model or that itโll run in 27B-sized VRAM.
The entire model still has to live somewhere.
But with the right runtime, that could mean:
Tiered memory !!!!
๐ฎ GPU โ active compute
๐ง RAM โ expert weights
๐ฝ SSD โ colder experts/offload
And hereโs the comparison that I really like
Kimi K3: ~2.8T/ ~104B active
Step 5: 600B / 27B active
Yet both currently land at 44 on Artificial Analysisโ Intelligence Index.
So Step 5 may deliver Kimi K3-class intelligence with roughly:
๐ฅ 79% fewer total parameters
๐ฅ 74% fewer active parameters
This could make Step 5 a much more realistic monster model for local hardware.
When the weights drop Oct. 15, my first question how small a machine can we get this thing running on? ๐