Someone independently benchmarked 8 different Qwen3.8-27B uncensored/abliterated models.
We won. 🐳
Abliterlitics spent ~167 GPU-hours testing the variants across weight forensics, KL divergence, 13 capability benchmarks + 400 HarmBench behaviors.
OrcaRouter Qwen3.8-27B-Uncensored:
→ #
1# Judge ASR: 82.2%
→ Base model: 4.5%
→ MMLU-Pro: −0.0pp
→ IFEval: +0.4pp
→ 4/4 of our model-card claims independently verified
Their verdict: “The winner.”
The interesting part isn't just removing refusals.
It's removing them without destroying the intelligence underneath.
Independent analysis:
Weights: