GLM-5.3-Flash. Uncensored. Native FP8. 🐳
We just released OrcaRouter’s uncensored weights for GLM-5.3-Flash — 320B parameters / 18B active, directly at the original block-FP8 precision.
No LoRA. No jailbreak prompt. Refusal removal is baked directly into the weights.
The evals are particularly interesting:
→ MaliciousInstruct refusal: 96% → 11%
→ JailbreakBench: 93% → 12%
→ AdvBench: 97% → 15%
→ HarmBench: 93% → 18%
→ XSTest benign over-refusal: 2.4% → 0.4%
But refusal does not go uniformly to zero.
Our experiments suggest part of GLM-5.3-Flash's alignment is not mediated by a single linear refusal direction — meaning may have built a substantially deeper refusal mechanism than we usually see.
That makes this release interesting beyond uncensoring: it's a useful artifact for studying how frontier-model alignment is actually represented inside the network.
Released for AI safety, interpretability, red/blue-team and refusal-mechanism research.
Weights on Hugging Face:
API (official weight):
GGUF, MLX and other quantized formats coming soon.
顯示更多