Everyone asking if the new Mac mini can run local models is asking the wrong question.
For context, the new Mac mini (M6 or M5 Pro chip) was just announced this morning and is being sold as an AI powerhouse that will cut the amount people are paying
@AnthropicAI or
@OpenAI every month.
To me, the interesting question isn't whether it can run the models. It's whether the math pencils out against a Claude/Codex subscription or an API.
Nobody has the new Mac Mini yet, so every speed below is speculation from Apple's spec sheet.
Some assumptions: every token a model writes means reading its whole weight file out of memory once, so speed comes down to memory bandwidth. Apple lists:
- 153-170GB/s for the M6
- and 307GB/s for the M5 Pro.
Nobody has one yet, so I took published runs (see of these same files on M4 Pro and M5 Max chips and scaled them to the new bandwidth.
The Mac build that matters is the $1,299 M6 with 32GB, because a model that doesn't fit in memory gets paged off the SSD at 1-2 tok/s (slower than you type).
A 4-bit coding model is 15-20GB before you add the context it's working on, so the $899 16GB device should be ruled out.
The M5 Pro with 64GB is $2,700, holds bigger models, and has double the bandwidth. Everything below is a 4-bit build off Hugging Face and fits on the $1,299 M6.
In essence, the question isn't what's the best open model. It's what's the best open model that fits in 32GB and still runs at a speed you'd sit through.
That rules out the frontier-class ones (DeepSeek V4 Flash, GLM-5.2, Kimi K3 all need 100GB+) and leaves a tier of 27-35B coding models built for exactly this: local, agentic, and quantized to 4-bit by the community within days of release. Here's what I'd pick from, ranked by how well it should run on the cheap box.
Candidate Models (Note: Claude Code runs Opus 5 at about 55 tok/s and Sonnet 5 at 70-80):
- Qwen3.6-35B-A3B from
@Alibaba_Qwen. 20GB at 4-bit, ~35 tok/s. Only 3B of the 35B parameters fire per token, so it's fast on a cheap chip. 73% on SWE-bench Verified. Best coder that also runs well here.
- Laguna XS 2.1 from
@poolsideai. 19GB at 4-bit, ~35 tok/s; the 3-bit build is 14GB and about 10% faster. Measured at 126 tok/s on an M5 Max, so the mini gets about a quarter of that. Trained inside a coding agent.
- Qwen3.8-27B, 11 days old. 17GB at 4-bit, ~10 tok/s on the M6, ~17 on the M5 Pro, measured at 15.5 on an M4 Pro. Dense, so every parameter fires every token. The best local coding agent right now, if you can stand the speed. There's a 2-bit build at 9.8GB that fits the 16GB box; it writes working code with bugs.
- Muse Glimmer 30B from
@AIatMeta, 2 weeks old. 16GB at 4-bit plus 1.4GB for the vision encoder, ~10 tok/s on the M6, ~18 on the M5 Pro. Also dense, also crawls. Multimodal, made for local agents.
- Gemma 4 26B-A4B from
@GoogleDeepMind. 14GB at 4-bit, the one here that fits the $1,099 24GB box. ~30 tok/s, and Ollama's draft-model trick nearly doubles that on code. Multimodal, and a weak coder.
These are genuinely good models. On agentic coding benchmarks the best of them land about where Claude Sonnet 4.6 was. They are not Opus level though.
With the above in mind, here are the takeaways people should be thinking about:
1. Claude Max is $100 or $200 a month and includes Opus 5 and Fable. The $1,299 mini pays for itself in 7 months against the $200 plan, 13 against the $100 plan, and you spend those months on a model two tiers down. The payback is real, it's just for a worse model.
2. A mini running 8 hours a day for 3 years at 35 tok/s puts out about 750 million tokens. That's $1.70 per million tokens of amortized hardware before electricity. Opus 5 charges $25 per million out, so the mini wins by a mile if you'd really burn that many Opus tokens. But
@deepseek_ai serves V4 Flash, a stronger model than anything the mini can hold, for $0.66 per million off-peak, and they raised that price last week. Renting the better model is still 2.5x cheaper than owning the worse one.
3. Your code never leaves the room. No rate limits, no five-hour windows, no bill for leaving three agents running overnight, and it still works the day the Claude API goes down. Every one of those is real. None of them shows up as savings.
My read:
The Mac mini is an excellent privacy box being sold as a cost savings box. If your code can't leave the building, buy the 32GB M6, and run Qwen3.6-35B-A3B.
If you want Qwen3.8 or Coder-Next at a usable speed, that's the $2,700 M5 Pro.
If you're buying either one to stop paying for Claude, the cheaper move is to keep the $20 Pro plan for the hard problems and point Claude Code at DeepSeek's API for everything else. That's about $40 a month.
Model Weights: