Qwen3.8-27B Unsloth GGUF is now the #1# most-liked GGUF of all time!
The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥
GGUF:
Guide:
Qwen3.8-27B Unsloth GGUF is now the #2# trending model on Hugging Face with 2.7M downloads! 💗
Unsloth also reached #3# trending on GitHub!
Thanks so much for the love!
Model:
GitHub:
Qwen3.8-27B GGUF has reached 1,000 likes in less than 24 hours! ❤️
It's now the #3# trending model on Hugging Face with 1M overall downloads.
Run on 17GB RAM/VRAM setups via Unsloth!
Model:
GitHub:
We released experimental MTP Qwen3.6 Unsloth GGUFs!
Qwen3.6 27B MTP now runs at 140 tokens/s. Qwen3.6 35B-A3B MTP gets 220 tokens/s generation on a single GPU.
Qwen3.6 27B and 35B-A3B have >1.4x speed-up over the original GGUFs without any change in accuracy.
Guide + GGUFs + Benchmarks:
In terms of average speedup, we see a 1.4x for dense models at draft tokens = 2 and for the MoE around 1.15 to 1.2x.
We do not recommend more than 2 draft tokens because the acceptance rate drops precipitously from 83% to 50% with 4 draft tokens, and the forward passes for MTP become less beneficial.
Use `--spec-type mtp --spec-draft-n-max 2`
Thanks to Aman for
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️
MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss.
Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s.
GGUFs + Guide:
Standup Pulse: an open-source project running async Slack standups using Gemma 4 26B-A4B (GGUF via llama.cpp) on an Apple M5 Max. It pairs Mastra for typed tool selection with CopilotKit Channels for restrained Slack Block Kit (Slack's native layout format) interactions while keeping all model inference, standup records, and traces local in SQLite.
This architecture is a great demonstration of how to connect Slack to a project without exposing the local server!
🔗 Blog:
🔗 Repo:
I’m evaluating a batch of quantized Qwen3.8 27B models (non-GGUF) on 20K+ prompts across a variety of tasks:
cyankiwi/Qwen3.8-27B-AWQ-INT4 — done
unsloth/Qwen3.8-27B-NVFP4 — done
EschaLabs/Qwen3.8-27B-Escha-W2 — done
ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ — running
Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw — pending (still figuring out whether it can support high concurrency; if not, it won't make it to the list)
I can include one more model. Any recommendations (must be non-GGUF)? Maybe NVIDIA's NVFP4?
I’m also evaluating all the models I made here:
So far, everything looks good except for the smallest model, whose evaluation is still running.
Some of my quantizations use 8-bit activations. While they don’t appear to hurt accuracy, I’m not seeing any significant speedup either (on Blackwell). I’ll run a few ablation experiments with 16-bit activations to get a better sense of the tradeoff.
These are non-agentic evaluations. For agentic evals, I’ll focus on two models: the fastest 4-bit version and the lowest-bpw model that preserves accuracy.
Once everything is done, I'll publish all the results, including latency and token efficiency.
Opus-level intelligence running fully local at home on dual RTX 4090s.
Qwen3.8-27B-GGUF:
- 80 tok/s
- Full 262k context
- MTP on
- Only 34 GB of 48 GB VRAM used
- Outperforms Claude Opus max on SWE-Pro: 61.7 vs Claude Opus 4.6’s 53.4 GPQA: 89.2 vs 91.3
No need API bills, Or no rate limits.
Local frontier-class models are no longer theoretical.
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗
Qwen3.8-27B GGUF is already Unsloth’s #1# most-downloaded model ever.
Thanks for all your support!