đ New Blog: Running DeepSeek-V4-Flash and Kimi-K3 on Consumer Hardware with SSD Expert Pack
These models are far too large for a typical PC's memory. WiCi AI and the SGLang team built SSD Expert Pack: routed experts stay on an NVMe SSD, and the runtime loads only the experts the router selects into a GPU cache.
On one RTX 5090, 32 GB RAM, and a 2 TB SSD
đ¸ DeepSeek-V4-Flash MXFP4: 1.85â1.99 tokens/sec decode
đ¸ Kimi-K3 community Q2_K (text-only): ~0.29 tokens/sec decode
Thanks to the WiCi AI team for the collaboration!
Read the full blog below đ