Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about.
đĄ I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick)
My llama.cpp recipe đ
đ§ Qwen3.5-4B Q4_K_M GGUF
âī¸ Ryzen 5 7540U â 6C/12T
đ§ĩ --threads 9
đ§ĩ --threads-batch 12
⥠--prio 2
đ --poll 50
đĻ --batch-size 2048
đĻ --ubatch-size 512
đ --flash-attn on
đ§ KV cache: q4_0 / q4_0
đ§ --repack
đž --mmap
đ¤ --parallel 1
đĢ --device none
đĢ --gpu-layers 0
đĢ KV/op GPU offload
đĢ MTP OFF
Interesting result đ MTP=3 was slower (~10 tps in benchmarks).
Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s.
That's about 10% faster just from tuning llama.cpp â on a basic laptop CPU.