🤯 A 600B parameter model is going open-weight Oct. 15…
…but only 27B parameters are active per token. 👈👀
Here is Step 5 Preview (this model is exciting) and here is why …
Stats 👇
🧠 600B total parameters
⚡ 27B active/token
📚 1M context
👁️ vision + video
🤖 built for long-running agents
🔓 open weights Oct. 15
And that 27B-active number makes this really interesting for Local AI right?
At roughly 4-bit, 600B parameters would theoretically be ~300GB of weights.
Real-world quantization + overhead means I’d expect something more like ~320–350GB before accounting for KV cache and other runtime memory.
So we’re potentially looking at
💾 ~384GB-class memory → Q4 territory
💾 ~256GB-class memory → aggressive Q3 territory
And because this is MoE, only 27B parameters participate in each token’s compute.
⚠️ That does NOT mean this is a 27B model or that it’ll run in 27B-sized VRAM.
The entire model still has to live somewhere.
But with the right runtime, that could mean:
Tiered memory !!!!
🎮 GPU → active compute
🧠 RAM → expert weights
💽 SSD → colder experts/offload
And here’s the comparison that I really like
Kimi K3: ~2.8T/ ~104B active
Step 5: 600B / 27B active
Yet both currently land at 44 on Artificial Analysis’ Intelligence Index.
So Step 5 may deliver Kimi K3-class intelligence with roughly:
🔥 79% fewer total parameters
🔥 74% fewer active parameters
This could make Step 5 a much more realistic monster model for local hardware.
When the weights drop Oct. 15, my first question how small a machine can we get this thing running on? 👀