Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️
GGUFs can reach 170 tokens/s on a RTX PRO 6000.
MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change.
GGUFs:
Guide:
Qwen3.8-Flash can now be run locally! 🔥
The 125B MoE model outperforms Claude-Opus-4.6 (Max).
Run on 75GB RAM via Unsloth GGUFs.
Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds.
Guide:
GGUF: