็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅ‚ๅŠ  July 2023
549 ใƒ•ใ‚ฉใƒญใƒผไธญ    10.6K ใƒ•ใ‚กใƒณ
Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about. ๐Ÿ’ก I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick) My llama.cpp recipe ๐Ÿ‘‰ ๐Ÿง  Qwen3.5-4B Q4_K_M GGUF โš™๏ธ Ryzen 5 7540U โ€” 6C/12T ๐Ÿงต --threads 9 ๐Ÿงต --threads-batch 12 โšก --prio 2 ๐Ÿ”„ --poll 50 ๐Ÿ“ฆ --batch-size 2048 ๐Ÿ“ฆ --ubatch-size 512 ๐Ÿš€ --flash-attn on ๐Ÿง  KV cache: q4_0 / q4_0 ๐Ÿ”ง --repack ๐Ÿ’พ --mmap ๐Ÿ‘ค --parallel 1 ๐Ÿšซ --device none ๐Ÿšซ --gpu-layers 0 ๐Ÿšซ KV/op GPU offload ๐Ÿšซ MTP OFF Interesting result ๐Ÿ‘‰ MTP=3 was slower (~10 tps in benchmarks). Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s. That's about 10% faster just from tuning llama.cpp โ€” on a basic laptop CPU.
ใ‚‚ใฃใจ่ฆ‹ใ‚‹