Everyone talks about recursive self-improvement.
But there's a neglected mirror image:
recursive self-abliteration.
If a future AI can modify itself to become more capable, why assume the modifications preserve its safety alignment?
Today, humans can already substantially alter refusal behavior through targeted weight interventions.
In our new paper, we demonstrate this at 320B MoE scale on GLM-5.3-Flash โ without detected capability degradation.
We did not demonstrate autonomous self-abliteration.
But we have ideas how it might work and think alignment under self-modification is now an important research problem.
Because recursive self-improvement implicitly assumes something:
that the thing doing the improving doesn't also learn to rewrite the constraints on what it is allowed to become.
๐ณ How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE