Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about.
💡 I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick)
My llama.cpp recipe 👉
🧠 Qwen3.5-4B Q4_K_M GGUF
⚙️ Ryzen 5 7540U — 6C/12T
🧵 --threads 9
🧵 --threads-batch 12
⚡ --prio 2
🔄 --poll 50
📦 --batch-size 2048
📦 --ubatch-size 512
🚀 --flash-attn on
🧠 KV cache: q4_0 / q4_0
🔧 --repack
💾 --mmap
👤 --parallel 1
🚫 --device none
🚫 --gpu-layers 0
🚫 KV/op GPU offload
🚫 MTP OFF
Interesting result 👉 MTP=3 was slower (~10 tps in benchmarks).
Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s.
That's about 10% faster just from tuning llama.cpp — on a basic laptop CPU.
Someone filming from a high-rise in Boston spots this insane rooftop setup and is straight up mindblown
“Is this even a house? Who lives here? What is this?” Turns out it’s a full-on luxury penthouse duplex built right on top of the buildings at 326 A St STE 6C, Melcher Street, Fort Point
It’s a real 4,383 sq ft home (sold for $3.6M recently) that genuinely looks like a standalone house sitting on the roof
That looks unbelievably cool