註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Daniel Kokotajlo
@DKokotajlo
加入 November 2023
288 正在關注    52.5K 粉絲
Another good essay from an OpenAI capabilities researcher I know (well, now former). I think focusing on control and abandoning alignment for now is is a reasonable strategy, but implementing it effectively will require an agreement not to RSI to ASI at max or close to max speed. Figuring out control is going to take time and money. Also, it only works up to a point, I think. In AI 2040 we talk about this.
顯示更多
We don’t know how to reliably align frontier AI models. Alignment may get harder, not easier, as models become more capable. But maybe AI control can scale faster than dangerous capabilities. Let’s build a Great AI Firewall - without building totalitarian control of society. *Alignment is Dead (for now), Long Live Control!* We do not know how to specify or guarantee robust alignment of sufficiently capable models. Instead we should treat these models as untrusted. Optimizing models against imperfect reward signals may increase the propensity for reward hacking, deception or scheming. Models can also be jailbroken or fine-tuned away from their safeguards. The central bet of AI control is that our ability to constrain an AI’s access to the world may scale faster than our ability to understand or align what is happening inside it. Where possible, make verifying an answer easier than generating it. Formally verified software and hardware are unusually attractive: let AI produce the implementation and a machine-checkable proof that it satisfies the specification. This does not solve specification: the specification and verification stack remain part of the trusted base. Every consequential external action should pass through an independent control layer: code execution, network access, calling APIs, accessing physical systems - even communication with users. Optimize control monitors for extremely high recall of dangerous actions, even at the cost of rejecting many safe actions. Suspicious actions should be rejected or escalated. Every major increase in dangerous capability should require evidence that control still works. Chain-of-thought monitoring is valuable while it works, but it should not be foundational. Models can learn to obscure their reasoning and may increasingly reason without fully verbalizing it. Open weights shift much of control from the model developer to whoever deploys the model. Technical control is only useful if dangerous deployments cannot cheaply route around it. We should approach AI control like cybersecurity. The goal is defense in depth: make successful catastrophic attacks sufficiently difficult, expensive and rare. Call this technical and institutional architecture the Great AI Firewall: the boundary between untrusted frontier intelligence and consequential real-world power. The name is deliberately provocative. China has substantial experience building large-scale technical control infrastructure. That may create some common ground for international coordination. But the analogy is also a warning. AI control must not become control of society. The goal is to constrain dangerous machine capabilities - not human speech, actions, or ordinary access to information. Controls should scale with capability and risk. Ordinary models should face ordinary constraints. More consequential capabilities justify stronger controls. The objective is the minimum control necessary to keep catastrophic risk acceptably low - not maximum control for its own sake. Firewall the AI, not society.
顯示更多
0
14
321
20
轉發到社區