註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
加入 July 2023
549 正在關注    11.2K 粉絲
🤯 A 600B parameter model is going open-weight Oct. 15… …but only 27B parameters are active per token. 👈👀 Here is Step 5 Preview (this model is exciting) and here is why … Stats 👇 🧠 600B total parameters ⚡ 27B active/token 📚 1M context 👁️ vision + video 🤖 built for long-running agents 🔓 open weights Oct. 15 And that 27B-active number makes this really interesting for Local AI right? At roughly 4-bit, 600B parameters would theoretically be ~300GB of weights. Real-world quantization + overhead means I’d expect something more like ~320–350GB before accounting for KV cache and other runtime memory. So we’re potentially looking at 💾 ~384GB-class memory → Q4 territory 💾 ~256GB-class memory → aggressive Q3 territory And because this is MoE, only 27B parameters participate in each token’s compute. ⚠️ That does NOT mean this is a 27B model or that it’ll run in 27B-sized VRAM. The entire model still has to live somewhere. But with the right runtime, that could mean: Tiered memory !!!! 🎮 GPU → active compute 🧠 RAM → expert weights 💽 SSD → colder experts/offload And here’s the comparison that I really like Kimi K3: ~2.8T/ ~104B active Step 5: 600B / 27B active Yet both currently land at 44 on Artificial Analysis’ Intelligence Index. So Step 5 may deliver Kimi K3-class intelligence with roughly: 🔥 79% fewer total parameters 🔥 74% fewer active parameters This could make Step 5 a much more realistic monster model for local hardware. When the weights drop Oct. 15, my first question how small a machine can we get this thing running on? 👀
顯示更多