You don’t need 128GB of unified memory to run Qwen 3.8 Next Flash & DeepSeek V4 Flash.
Run Qwen 3.8 Next Flash on Mac 20GB RAM
it Wallm streams routed experts off the SSD instead of loading the full checkpoint into unified memory.
- feeds on Codex, 11K input: 60 tok/s prefill, 8 tok/s decode.
- TTFT is the pain: 17–76 s depending on prompt length
- Peak RAM: 23 to 36 GiB as context grows
This is not MLX loading a 2-bit quant into 96GB+. Custom Swift/Metal runtime, Experts live on NAND until the router asks for them.