Register and share your invite link to earn from video plays and referrals.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
Joined July 2023
549 Following    10.6K Followers
Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about. 💡 I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick) My llama.cpp recipe 👉 🧠 Qwen3.5-4B Q4_K_M GGUF âš™ī¸ Ryzen 5 7540U — 6C/12T đŸ§ĩ --threads 9 đŸ§ĩ --threads-batch 12 ⚡ --prio 2 🔄 --poll 50 đŸ“Ļ --batch-size 2048 đŸ“Ļ --ubatch-size 512 🚀 --flash-attn on 🧠 KV cache: q4_0 / q4_0 🔧 --repack 💾 --mmap 👤 --parallel 1 đŸšĢ --device none đŸšĢ --gpu-layers 0 đŸšĢ KV/op GPU offload đŸšĢ MTP OFF Interesting result 👉 MTP=3 was slower (~10 tps in benchmarks). Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s. That's about 10% faster just from tuning llama.cpp — on a basic laptop CPU.
Show more