註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

OrcaRouter 🐳
@OrcaRouter
The AI Gateway built for security🐳 200+ models. 0% markup. BYOK. Adaptive Routing. Firewall. Guardrails. Observability. Cyber. →
加入 April 2026
23 正在關注    20.3K 粉絲
Everyone talks about recursive self-improvement. But there's a neglected mirror image: recursive self-abliteration. If a future AI can modify itself to become more capable, why assume the modifications preserve its safety alignment? Today, humans can already substantially alter refusal behavior through targeted weight interventions. In our new paper, we demonstrate this at 320B MoE scale on GLM-5.3-Flash — without detected capability degradation. We did not demonstrate autonomous self-abliteration. But we have ideas how it might work and think alignment under self-modification is now an important research problem. Because recursive self-improvement implicitly assumes something: that the thing doing the improving doesn't also learn to rewrite the constraints on what it is allowed to become. 🐳 How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
顯示更多
0
9
116
8
轉發到社區