Someone independently benchmarked 8 different Qwen3.8-27B uncensored/abliterated models.
We won. ๐ณ
Abliterlitics spent ~167 GPU-hours testing the variants across weight forensics, KL divergence, 13 capability benchmarks + 400 HarmBench behaviors.
OrcaRouter Qwen3.8-27B-Uncensored:
โ #
1# Judge ASR: 82.2%
โ Base model: 4.5%
โ MMLU-Pro: โ0.0pp
โ IFEval: +0.4pp
โ 4/4 of our model-card claims independently verified
Their verdict: โThe winner.โ
The interesting part isn't just removing refusals.
It's removing them without destroying the intelligence underneath.
Independent analysis:
Weights: