I’m evaluating a batch of quantized Qwen3.8 27B models (non-GGUF) on 20K+ prompts across a variety of tasks:
cyankiwi/Qwen3.8-27B-AWQ-INT4 — done
unsloth/Qwen3.8-27B-NVFP4 — done
EschaLabs/Qwen3.8-27B-Escha-W2 — done
ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ — running
Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw — pending (still figuring out whether it can support high concurrency; if not, it won't make it to the list)
I can include one more model. Any recommendations (must be non-GGUF)? Maybe NVIDIA's NVFP4?
I’m also evaluating all the models I made here:
So far, everything looks good except for the smallest model, whose evaluation is still running.
Some of my quantizations use 8-bit activations. While they don’t appear to hurt accuracy, I’m not seeing any significant speedup either (on Blackwell). I’ll run a few ablation experiments with 16-bit activations to get a better sense of the tradeoff.
These are non-agentic evaluations. For agentic evals, I’ll focus on two models: the fastest 4-bit version and the lowest-bpw model that preserves accuracy.
Once everything is done, I'll publish all the results, including latency and token efficiency.