Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about.
💡 I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick)
My llama.cpp recipe 👉
🧠 Qwen3.5-4B Q4_K_M GGUF
⚙️ Ryzen 5 7540U — 6C/12T
🧵 --threads 9
🧵 --threads-batch 12
⚡ --prio 2
🔄 --poll 50
📦 --batch-size 2048
📦 --ubatch-size 512
🚀 --flash-attn on
🧠 KV cache: q4_0 / q4_0
🔧 --repack
💾 --mmap
👤 --parallel 1
🚫 --device none
🚫 --gpu-layers 0
🚫 KV/op GPU offload
🚫 MTP OFF
Interesting result 👉 MTP=3 was slower (~10 tps in benchmarks).
Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s.
That's about 10% faster just from tuning llama.cpp — on a basic laptop CPU.
Qwen3.8-27B Uncensored Cyber model you can run
15 GB locally
- 528 real tool calls
- 262K context.
- Aggressive abliteration to Cyber peel
- Imatrix from real agent logs not wiki text.
- Real tool-call calibration.
-
Qwen3.8-27B on 1500$ oh hardware.
This is the realistic local AI experience for most people, for tons of reasons including complexity, costs, and over abundance of configs.
Still, usable. GLM-4 was about 30 tok/s at peak
The future is bright.
Goodnight friends.
Qwen3.8-Flash API is live on QwenCloud. 262K native context, extensible to 1M, and priced to scale.
About Price⬇️
Input: $0.15 / 1M tokens
Output: $0.47 / 1M tokens
Cache hit: $0.016 / 1M tokens
Get your API key and start building:
#QwenCloud# #Qwen#
Qwen3.8-Flash-Next on a single RTX PRO 6000 👀
Quant: Unsloth AI UD-Q4_K_XL
Model size: 111 GB
GPU: RTX PRO 6000, 96 GB GDDR7
Fully on GPU with full context.
Benchmark context tested up to ~253K
Results:
- 40.3 t/s average decode across 8K–253K
- 63 t/s peak decode @ 8K
- 17.6 t/s decode @ ~253K
- 1,657 t/s peak prefill @ 8K
All results were run without MTP or DFLASH.
Qwen3.8 Flash from @Alibaba_Qwen is now live on OpenRouter.
A multimodal reasoning model for coding assistants, agentic workflows, visual understanding, codebase and document analysis, desktop interaction, charts, and long video.
Qwen3.8-27B Q3_K_XL is downloading right now. Going to see how much context I can squeeze out of it on a 20GB RX 7900 XT without killing the speed. This could be surprisingly usable or turn into a very long evening lol.
Qwen3.8-27B just got stupid fast.
Moved it to ½ RTX PRO 6000 and the prefill/decode speed is absurd compared to L4.
latest optimizations with SGLang completely changed the speed profile.
If you’re running this on RTX PRO 6000 or DGX Spark update the repo right now, It’s worth it.
Qwen3.8-2.4T-A95B, compressed two ways at once: 25% of the experts pruned with REAP, and the rest quantized to NVFP4.
Even with a quarter of the experts gone and 4-bit weights, GPQA Diamond holds at 91.5 vs 92.6 for the full-precision base. ~99% recovery.
Serve on @vllm_project: