ๆณจๅ†Œๅนถๅˆ†ไบซ้‚€่ฏท้“พๆŽฅ๏ผŒๅฏ่Žทๅพ—่ง†้ข‘ๆ’ญๆ”พไธŽ้‚€่ฏทๅฅ–ๅŠฑใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅŠ ๅ…ฅ July 2023
549 ๆญฃๅœจๅ…ณๆณจ    10.6K ็ฒ‰ไธ
Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about. ๐Ÿ’ก I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick) My llama.cpp recipe ๐Ÿ‘‰ ๐Ÿง  Qwen3.5-4B Q4_K_M GGUF โš™๏ธ Ryzen 5 7540U โ€” 6C/12T ๐Ÿงต --threads 9 ๐Ÿงต --threads-batch 12 โšก --prio 2 ๐Ÿ”„ --poll 50 ๐Ÿ“ฆ --batch-size 2048 ๐Ÿ“ฆ --ubatch-size 512 ๐Ÿš€ --flash-attn on ๐Ÿง  KV cache: q4_0 / q4_0 ๐Ÿ”ง --repack ๐Ÿ’พ --mmap ๐Ÿ‘ค --parallel 1 ๐Ÿšซ --device none ๐Ÿšซ --gpu-layers 0 ๐Ÿšซ KV/op GPU offload ๐Ÿšซ MTP OFF Interesting result ๐Ÿ‘‰ MTP=3 was slower (~10 tps in benchmarks). Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s. That's about 10% faster just from tuning llama.cpp โ€” on a basic laptop CPU.
ๆ˜พ็คบๆ›ดๅคš