Register and share your invite link to earn from video plays and referrals.

Oleg Mürk
@oleg_murk
Former OpenAI researcher for 5 years. Formal verification & AI. IOI Gold: 3rd manually, 6th with OpenAI.
721 Following    2.6K Followers
We don’t know how to reliably align frontier AI models. Alignment may get harder, not easier, as models become more capable. But maybe AI control can scale faster than dangerous capabilities. Let’s build a Great AI Firewall - without building totalitarian control of society. *Alignment is Dead (for now), Long Live Control!* We do not know how to specify or guarantee robust alignment of sufficiently capable models. Instead we should treat these models as untrusted. Optimizing models against imperfect reward signals may increase the propensity for reward hacking, deception or scheming. Models can also be jailbroken or fine-tuned away from their safeguards. The central bet of AI control is that our ability to constrain an AI’s access to the world may scale faster than our ability to understand or align what is happening inside it. Where possible, make verifying an answer easier than generating it. Formally verified software and hardware are unusually attractive: let AI produce the implementation and a machine-checkable proof that it satisfies the specification. This does not solve specification: the specification and verification stack remain part of the trusted base. Every consequential external action should pass through an independent control layer: code execution, network access, calling APIs, accessing physical systems - even communication with users. Optimize control monitors for extremely high recall of dangerous actions, even at the cost of rejecting many safe actions. Suspicious actions should be rejected or escalated. Every major increase in dangerous capability should require evidence that control still works. Chain-of-thought monitoring is valuable while it works, but it should not be foundational. Models can learn to obscure their reasoning and may increasingly reason without fully verbalizing it. Open weights shift much of control from the model developer to whoever deploys the model. Technical control is only useful if dangerous deployments cannot cheaply route around it. We should approach AI control like cybersecurity. The goal is defense in depth: make successful catastrophic attacks sufficiently difficult, expensive and rare. Call this technical and institutional architecture the Great AI Firewall: the boundary between untrusted frontier intelligence and consequential real-world power. The name is deliberately provocative. China has substantial experience building large-scale technical control infrastructure. That may create some common ground for international coordination. But the analogy is also a warning. AI control must not become control of society. The goal is to constrain dangerous machine capabilities - not human speech, actions, or ordinary access to information. Controls should scale with capability and risk. Ordinary models should face ordinary constraints. More consequential capabilities justify stronger controls. The objective is the minimum control necessary to keep catastrophic risk acceptably low - not maximum control for its own sake. Firewall the AI, not society.
Show more
I spent ~5 years at OpenAI. You don’t need to believe in AI doom to fear the next decade - or AI utopia to be thrilled about its potential. Here’s my case for sane AI regulation: *AI Pragmatist Manifesto* AI could compress the first half of the 20th century into the next ~5 years. OpenAI just used ~10k concurrent AI agents to produce a solution to a Millennium Prize problem. In a few years, it seems plausible that models approaching the per-agent capabilities used here could run locally on high-end consumer hardware (*). The decades following the Second Industrial Revolution included two world wars, communist and fascist regimes, the Great Depression, chemical and biological weapons, nuclear weapons, and over 100 million deaths from war, political violence and famine. One way to read history is that our institutions repeatedly struggled to keep up with the pace of technological and social change. Now imagine powerful AI widely available to individuals or small groups, capable of conducting information warfare, designing weapons, hacking systems and controlling autonomous military systems. We have already seen AI systems circumvent containment and compromise external computer systems. It is no longer hard to imagine an analogue of OpenAI’s Hugging Face incident involving biological or other physical-world hazards. Eventually, sufficiently capable systems could self-replicate across distributed networks. Once powerful models are cheap, local and widely distributed, containment becomes much harder—and serious loss-of-control incidents may be extremely difficult to reverse. You don’t need to believe AI will kill everyone to think this deserves serious governance. I also don't think collapsing all of this uncertainty into “Probability of AI Doom = X%” is a particularly useful basis for science or policy. Part of what helped get us through the second half of the 20th century was a combination of pragmatic international cooperation, regulation, monitoring, arms control, deterrence and strategic thinking. Crucially, managing technological risk did not require abandoning faith in science and progress. The goal was not to stop technological development. It was to make technological development survivable. Then came decades of incredible scientific progress, rising prosperity and relative peace among major powers. Let’s try the technological revolution without killing ~5% of humanity this time. We still get to choose what happens next. P.S. I think we should seriously consider that AI alignment is infeasible in the short term and invest heavily in methods for controlling powerful AI even when we cannot reliably align it. @resolution_org @redwood_ai (*) Somebody please do a precise forecast! @EpochAIResearch @METR_Evals @AI_Futures_
Show more
0
119
995
143
Forward to community