登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

OrcaRouter 🐳
@OrcaRouter
The AI Gateway built for security🐳 200+ models. 0% markup. BYOK. Adaptive Routing. Firewall. Guardrails. Observability. Cyber. →
参加 April 2026
23 フォロー中    20.3K ファン
Everyone talks about recursive self-improvement. But there's a neglected mirror image: recursive self-abliteration. If a future AI can modify itself to become more capable, why assume the modifications preserve its safety alignment? Today, humans can already substantially alter refusal behavior through targeted weight interventions. In our new paper, we demonstrate this at 320B MoE scale on GLM-5.3-Flash — without detected capability degradation. We did not demonstrate autonomous self-abliteration. But we have ideas how it might work and think alignment under self-modification is now an important research problem. Because recursive self-improvement implicitly assumes something: that the thing doing the improving doesn't also learn to rewrite the constraints on what it is allowed to become. 🐳 How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
もっと見る