This is a crazy, weird Qwen3.8-27B โuncensoredโ release.
It isn't because it's uncensored, but because the author measured what the guardrail surgery cost to get it uncensored and then preserved MTP.
Jonathan Coletti took Qwen3.8-27B and used Heretic model surgery to reduce its refusal behavior.
Measured on ~100 held-out test prompts ...
๐ Base Qwen3.8-27B โ w/98 refusals
๐ Modified model โ w/only 12 refusals
Meanwhile, the model's capability barely changed ...
๐ง MMLU 83.4 โ 83.3
๐ฏ HellaSwag 82.8 โ 82.9
๐ Mean across 4 tests โ ~0.5 point drop
Abliteration initially dropped Qwen's MTP head and the author then grafted all 15 MTP tensors back from the original Qwen3.8-27B and then verified that the MTP block actually survived GGUF quantization.
So the modified model you still get ...
๐ง 27B dense Qwen
๐๏ธ multimodal/vision
๐ 262K context architecture
๐ฎ native MTP speculative decoding
๐ฆ llama.cpp GGUF
And mradermacher already has iMatrix quantizations
๐ฆ IQ3_S โ 12.7GB
๐ฆ IQ3_M โ 12.9GB
๐ฆ Q3_K_M โ 13.6GB ๐
๐ฆ IQ4_XS โ 15.4GB
๐ฆ Q4_K_M โ 16.9GB
That makes ~13GB versions especially interesting for 16GB GPU testing. ๐ฅ
โ ๏ธ File size โ VRAM requirement, so don't assume the 15.4GB IQ4 automatically fits comfortably in 16GB with useful context.
Also, I have found that โuncensoredโ model variants don't mean capability is proven unchanged.
For example, their published comparison doesn't yet cover coding, math, multilingual, or vision quality. ๐