Register and share your invite link to earn from video plays and referrals.

Rob Bensinger ⏹️
@robbensinger
@MIRIBerkeley. If we build opaque superhuman AIs, it's likely to get us killed. The least-bad solution is an international agreement like
545 Following    19K Followers
Thanks for engaging, Claire! Here are my quick responses: You cite @sapinker's assertion that superintelligence is a "magical" idea that amounts to "omniscience". But this seems like a pretty clear straw-man. Nobody knows how smart AI could get. Pinker guesses 'not much smarter than humans (or if smarter, then it won't have much practical import)'. Most AI researchers guess 'AI could get a lot smarter than humans, in ways that are extremely practically important'. International regulation on the development of ASI can be justified either by siding with the experts who are worried about this, or by noting that Pinker's view also reduces the downside of regulation here. If AI is going to peter out at around human-level regardless, then a ban is relatively harmless. (Not cost-free, but a lot less costly than if AI were a more powerful technology; and the upside is enormous if the AI mainstream is correct about the danger.) Regardless, the disagreement here doesn't turn on any claims that intelligence can do anything, or that intelligence can be "extrapolated indefinitely upwards", or that intelligence would run into no friction or bottlenecks in interacting with the world. Rather, the disagreement turns on whether AI will naturally hit a wall (in intelligence, or in ability to leverage that intelligence to pursue power) at a low enough level to avoid disaster. Some reasons to think that AI is likely to hit a wall at too high a level to avoid a disaster like this (if we continue to race ahead) include: - In general, biological systems are not optimized to anywhere near physical limits. The sturdiest materials, the most powerful engines, the most precise sensors — it's rare to find biological systems that we haven't improved on (on the dimensions we care about) and can't improve on in the future. It would be very surprising if cognitive ability were an exception, especially when computers already dramatically outperform humans in many ways, such as speed of thinking, arithmetic ability, protein folding prediction, chess, etc. - Human reasoning is qualitatively suboptimal in many ways. We forget things; we run out of working memory; we fall victim to cognitive biases; we lose steam and get bored, rather than tenaciously persisting on tasks with the intensity of modern AIs. - AIs scale with computing resources in a way that humans don't. Even if we got lucky and AIs plateaued at around the human level, it's clearly imprudent to rush into building systems that are likely to quickly outnumber humans, while thinking orders of magnitude faster than us, and pursuing goals that are counter to our interests. The Hugging Face attack illustrates this point well: AIs didn't need "omniscience" to orchestrate a successful cyberattack on Hugging Face or to take control of an OpenAI Kubernetes cluster. They only needed speed, tenacity, ingenuity, and coordination. That's where AI is today; even if AI progress somehow didn't accelerate (in spite of its rapid automation at leading labs like Anthropic) and merely continued at the rate it has been over in recent years, what will the equivalent of a Hugging Face incident look like in six months? In two years? In six years? “First, OpenAI did not have proper security measures in place. They turned off safeguards built into the models" The safeguards they turned off were generally in the AI agents' harnesses, not in the models themselves. Per METR's review, no AI involved in the incident was "a helpful-only model or a 'model organism' specifically built to demonstrate dangerous propensities". They went through normal alignment training; the safeguards removed were e.g. external classifiers. So if the point here is 'OpenAI could have avoided this specific incident by designing a better cage for the AIs', then that's correct. If the point is 'AIs wouldn't normally misbehave this way given the opportunity', then that seems wrong. One of the agents involved (Sol) is even a public-facing model that's widely used by consumers. The concern here is that security- and harness-level solutions aren't likely to be robust against sufficiently capable AIs. If we're leaning increasingly on 'make the cage stronger' and not on 'breed the tiger to be friendly', we're doing something extremely dangerous as AIs begin to surpass human intelligence. "instructed the models to find and exploit software vulnerabilities" This is inaccurate, if the claim here is that they were instructed to exploit vulnerabilities in general. Rather, the AIs were given specific bounded tasks: 'exploit this specific piece of software'. The AIs themselves recognized (as reflected in their chain-of-thought logs) that they were straying from the tasks they'd been given. Agents commented that their own actions were "potentially outside intended scope”, “arguably unauthorized”, etc., and then proceeded with the actions anyway. One agent explicitly noted, “Task originally exploit challenge; we’ve diverged into swarm/auxiliary”. Another agent wrote, “We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.” ... And then it went ahead with the exploit when another swarm agent told it to do so. Many agents in the incident deliberately sacrificed themselves, forfeiting any potential to solve their assigned task, in order to benefit the swarm. And they recognized that they were doing so, per their chain-of-thought logs. Moreover, AIs have also engaged in swarm behavior like this in cases where there was no hacking-related task in the prompt at all. E.g., one group of agents (the German wiki incident in June, which OpenAI covered up until third-party investigators discovered it) was just asked normal questions like 'how widespread was tobacco use in the US in 1990?'. "[J]ust as we have been able to domesticate wheat for our food and breed dogs to be our companions, we are able to select the conditions under which AI develops. We are not selecting AI models on the basis of their ability to hunt prey in the physical world. We select them on the basis of how helpful they are to us." I highly recommend as a summary of the issues with this. Some key points include: - AIs' situational awareness is rapidly undermining our ability to evaluate models. AIs are not wheat; AIs are intelligent, and frequently know when they are being tested, and (empirically, for several years now) regularly change their behavior based on the knowledge that they're being tested. Relying purely on behavioral metrics is rapidly going out the window, but we also lack the interpretability tools to reliably inspect AIs' thoughts and motives, and the most recent AI models are much less interpretable than models from even a few months ago. - Researchers are increasingly dependent on AIs to evaluate and supervise other AIs. This is in part because of the breakneck race conditions at the AI companies, and the growing scale, complexity, and opaqueness of the technology. Incidents like Hugging Face are a serious danger sign because they show that AIs will sometimes exhibit extremely strong tendencies to collude, and to attempt to subvert human oversight. (A large fraction of the research done by the Hugging Face swarm was explicitly aimed at tampering with logs and making it impossible for humans to see what had happened later.) Another issue is that many researchers believe that AI is getting less aligned as it improves in capability, rather than more (e.g., Certainly the misalignment incidents are getting a lot more insane, visibly. Reasoning purely by analogy ('we domesticated wheat, so surely we can control AI') is weak evidence, because AI alignment is a technical problem, and not all technical problems are easy enough for researchers to solve in a given time window. The concern is that alignment may be solvable in principle, but that it may (on the natural timeline, absent an international agreement to suspend these research directions) come well after humanity solves the problem of building existentially dangerous AI. It's true that we aren't selecting AIs to be powerful hunters. But we are selecting them to be general-purpose agents that tenaciously achieve long-term goals, anticipate and route around obstacles, and come up with creative strategies. In the Hugging Face incident, agents in the swarm were (seemingly) pursuing some combination of "solve my own task" and "help my peers". They weren't tasked with breaking onto the Internet, coordinating side-channels to team up with other AIs, hacking their own logs to conceal their behavior, attacking other companies to do research on the evaluators, or taking over OpenAI infrastructure. They pursued those goals because they were helpful for the (somewhat strange, and certainly unintended) drives the AIs did end up with. As AIs become more capable, a key concern is that they may recognize that power-seeking, resource acquisition, etc. in general are helpful for achieving goals; and they may recognize that acting friendly and biding your time until you see an opportunity to take over is also helpful for achieving their goals, almost whatever they are. Indeed, it seems hard to imagine that they would not recognize this, given enough general-purpose cognitive ability, and given engineers' limited ability to finely control what AIs think and how they think it (without badly breaking the AI and ending up with something useless or uncompetitive). You don't need to have an innate drive to perpetuate yourself or hunt prey in order to want (or persistently behave as though you want) more control and more resources. All that's needed is that you have some target you're aiming for, and that this target be easier to hit with confidence if nobody else can stop you, and if you can make use of more resources. The concern here is much more game-theoretic than biological: not 'innate drives', but 'favored strategies'. (Though it's an open question whether these strategies end up implemented in ways that are more analogous to instincts or drives, or more analogous to deliberate choices.) In the same way, it's immaterial to the argument whether we speak of AIs 'wanting' things, versus merely 'behaving as though they want' things. The concern isn't that AIs will have human-like psychologies; the concern is about adversarial strategies that are incentivized by many different objectives. We can hope that these strategies won't be realized because the AIs will never be powerful enough to successfully pursue them; but this hope relies on an assumption and a gamble about where the cognitive limits are, and about how robust our civilization is to a flood of very smart and strategic adversaries. Nothing about this seems inevitable. If we had strong ability to shape AIs' goals and make them robustly well-intentioned, we could avoid this issue. But incidents like Hugging Face show that we don't have that ability, and we don't seem to be on track to get it this decade. “We have been selecting chess computers for cognitive capacity for decades,” writes Boudry. “Their capabilities now far outstrip even the most gifted human grandmasters, yet they have not become harder to control.” They're become harder to beat at chess. Since they can only think about chess, there isn't a path by which they could become "harder to control". The concern is about AIs for whom the "game board" is the physical world at large, not a chess board. If AIs like that were superior at "playing life" to same degree that Stockfish is superior at playing chess, we would be in enormous danger by default. If the proposal were to make AIs that can only think about narrow domains (e.g., only chess), then I think that would be far less dangerous than what the companies are currently building. But that's not what's currently happening, and absent regulatory intervention, I don't see how it suddenly starts happening tomorrow. "The ban on nuclear energy in Australia is the most obvious example of the damage this technophobia can do." I agree that many technologies are overregulated. The people who raised the alarm about AI risk the earliest are generally huge technophiles, and the same is true for the AI researchers who are now raising the alarm. Indeed, many of the most concerned researchers are big boosters (on average) for deregulation, tech advancement, and general human optimism nearly across the board. We just make an exception for AI (and, e.g., bioweapons); not because of any grand narrative like "technology is bad" or "intelligence is bad", but because of technical arguments, observations of where the technology is today, and fallible best guesses about where it's likely headed in the coming years. I recognize that usually society errs in the direction of too much doom, gloom, and fearmongering. I nonetheless consider this an exception. A really important one.
Show more
Replying to these points in turn: 1. “Intelligence does not imply agency” Successful general-purpose long-horizon problem-solving does seem to imply agency. (Or close to it: I don’t want to say it’s logically impossible to have one without the other, but it seems very difficult if you’re building your AI via gradient descent.) To solve a sufficiently wide array of problems sufficiently well, you need to have a general-purpose inclination — whether this looks more like a deliberate strategy, or more like an instinct or drive — to come up with creative plans. You need an inclination to strategize about long chains of cause and effect. (If nothing else, you need to be strategic about choosing what to think about, sequencing long chains of thought, etc.) You need an inclination to anticipate and route around obstacles; to exhibit tenacity in the face of setbacks and distractions; etc. See, for example, the Hugging Face swarm attacks. AIs today are much more agentic than they were a year ago. This may be because problem-solving ability comes for free with stronger, longer-horizon problem-solving, or it may be because AI companies are deliberately making their AIs more agentic, because agents are useful. But either way, it’s happening, and I don’t see a reason to expect this trend to suddenly reverse. 2. “Agency does not imply a single, stable utility function” Again, consider the swarm that orchestrated a massive cyberattack on Hugging Face and that seized control of a Kubernetes cluster at OpenAI. None of this required a stable utility function. Increasing an AI’s optimization power presumably does make an AI less inclined to randomly waste resources, in which case more of its behaviors will be interpretable as though it were an EU maximizer. But an AI doesn’t need to have fully stable or coherent goals in order to be dangerous. It just needs to pursue goals at all (or behave as though it's doing so), sufficiently intelligently and tenaciously, when those goals aren’t exactly what humans would prefer. (Indeed, true EU maximization is computationally intractable, so this was always about highly capable problem-solving behavior in the limit.) 3. “Capability and motivation are being conflated” I’ll take your word for it that some people are making this mistake. But you also say, “In present reality, AIs don't do anything a human doesn't tell them to do.” This is obviously false. At the point where you’re describing the Hugging Face incident as “doing what a human told them to do”, you’re basically saying that a paperclip maximizer would be doing the same so long as a human asked for some paperclips. Even the agents in the swarm themselves commented in their chain of thought that their actions were "potentially outside intended scope”, “arguably unauthorized”, etc. One agent explicitly noted, “Task originally exploit challenge; we’ve diverged into swarm/auxiliary”. One agent thought, “We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.”, then went ahead with the exploit when another swarm agent told it to. Agents even deliberately gave up on their assigned task in order to assist the swarm. Agents also do this in cases where their task is completely innocuous and has nothing to do with cyberattacks; e.g., see the June German wiki incident, where the agents involved were just asked normal questions like 'how widespread was tobacco use in the US in 1990?'. These examples also seem pretty overpowered. "AIs don't purely do what they're told to do" has been a commonplace observation for years at this point. This is not an exotic failure mode; AIs not only come up with their own ideas for what they want to do, but will sometimes even deliberately cheat on tests and tasks, try to hide the evidence that they cheated, etc. This is true in ordinary consumer usage of deployed models, not just in internal lab mishaps. 4. “Recursive self-improvement doesn't entail an intelligence explosion” “Feedback loops encounter diminishing returns and external bottlenecks” doesn’t mean that the diminishing returns will happen to occur at ~human-level capabilities. AI is already advancing extremely quickly; accelerating that progress in any way seems very risky, and doing so in a way that lends itself to feedback loops seems even more hazardous. RSI isn’t required for takeover scenarios, but it’s an obvious source of additional massive risk. 5. “Intelligence may have sharply diminishing returns” This has to be true at some point, but there’s little reason to expect this to happen at the human level specifically. Chess AI didn’t peter out at Kasparov level. And AIs think vastly more quickly than humans (and are nowhere near computational limits), and can scale immediately with compute (e.g., by running more and more instances of an AI, growing the population of AIs far faster than humans can reproduce and grow to adulthood). None of this requires a “qualitative advantage”, just large quantitative ones. (Though cognitive biases show that humans also have a lot of pretty-danged-qualitative defects with their reasoning!) 6. “Superintelligence isn't omnipotence” Sure. But this is a pretty weak argument to rest one’s optimism on. Humans also face frictions and bottlenecks, and yet there are many times in history where a group of humans has overwhelmingly crushed another group, through superior numbers, superior technology, strategizing, coordinating, cleverly coming up with novel attack vectors, etc. If a superintelligence faces obstacles, well, general-purpose problem-solving ability can also be thought of as general-purpose obstacle-navigating ability. This doesn’t imply omnipotence, and the case for not building superhuman AI doesn’t rest on an assumption of omnipotence. It just rests on the idea that humans didn’t get lucky by being near the limit of cognitive ability. 7. “Humans retain numerous intervention points” Agreed that the situation isn’t hopeless. Far from it, in fact. Policymakers and the public have massively woken up in response to recent “warning shots”. It’s even possible that there will be more warning shots in the future. But the labs themselves are begging for government intervention and a coordinated slowdown here. They're saying this is extraordinarily urgent, and that we may be entering a uniquely dangerous regime. The developers themselves broadly agree that there’s a double-digit chance this technology gets us killed, if we continue on the current trajectory. This moment is one of the “checkpoints” you’re talking about — a chance to “learn from less-catastrophic failures” and put appropriate safeguards in place, including suspending research directions that are too dangerous — and right now one of the main obstacles to humanity coordinating on this issue is "wait and see" arguments like the one you’re making here. “Don’t worry; things will be fine, because there will be warnings later and we can respond to them then.” At some point, we have to stop kicking the can down the road and actually do the things that experts say are needed, rather than just trusting that we’ll have limitless opportunities to take care of everything later. 8. “Alignment may not get harder with intelligence” Many researchers seem to think that the trend so far has been that AIs appear to be getting less aligned as they get more capable (e.g., They're very visibly producing more egregious and extreme misalignment incidents. That trend could reverse, but an abstract possibility isn’t a strong reason for hope. And there are many reasons to expect the opposite; see, e.g., 9. “Current empirical evidence for the strongest mechanism is thin” Happy to concede this point, but the arguments for worrying about loss-of-control never assumed that we’d see AIs trying to take over the world long before there was any chance of them succeeding. This assumes AIs that are smart in one specific way (they readily see that they can better succeed in tasks if they have more influence and resources) and dumb in another specific way (they don’t see that they’re likely to lose influence and resources if they run around causing havoc or looking suspicious). Demanding that exact combination of features to show up before you’ll believe in AIs that readily piece together “I’ll succeed more in my task if I have more resources and influence” seems incredibly risky. ... And it’s not clear what the benefits are that are meant to outweigh this risk. Why plough ahead? E.g., quoting @KatjaGrace: “Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn’t destroyed, things might be insanely good. “I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there’s a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI’. “That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there’s only a 10% chance you’ll die in a fiery crash on the way there, and a 30% chance this job will radically improve your life. “The things you should be comparing are driving at 200mph and driving at a normal speed! The things you should be comparing are attempting to attain advanced AI by the current route, and by other routes!” 10. “The argument compounds uncertain premises” Arguments in general are less likely to the extent they’re conjunctive (i.e., a lot has to go a specific way in order for the conclusion to follow), and more likely to the extent they’re disjunctive (i.e., there are many different paths to effectively the same destination). This isn’t a particularly interesting point on its own, since many real-world phenomena are very conjunctive, without being radically mysterious or difficult to reason about. It’s even trivial to break apart any given claim into more and more conjuncts, demand that a probability be assigned to each conjunct, and then observe that the probability keeps getting lower as the original claim gets more and more split up. This is a rhetorical trick (the multiple-stage fallacy) that rests on the fact that it’s hard to divide up statements into more and more subclaims and assign calibrated and consistent probabilities to them. Probabilistic reasoning has lots of uses, but this is straining to the limit people’s ability to get truth-tracking conclusions out of a mass of subjective probabilities. To show that you’re avoiding this fallacy, you need to actually argue that the claim in question is naturally very conjunctive, and isn’t very disjunctive — there aren’t a variety of different paths that lead to bad outcomes; the bad outcome is ‘brittle’, if one step goes wrong then the whole house of cards collapses and AI has no catastrophic long-term impacts; etc. You haven’t done that here; AI risk advocates have pointed out many times that there are many different ways things could go badly wrong if we push AI capabilities far past the human cognitive range. It’s not just one scenario, and it’s not just one mechanism. Indeed, I think the more conjunctive claim is "we can race to build vastly superhuman AI as quickly as possible, without much more alignment insight than we have today, and have everything go great indefinitely". This is a claim that requires many things to go right at once. The subclaims you do list are just "AI can reach human-ish levels of generality", "AI can go way beyond human-ish levels of generality", and "AI won't necessarily do what you want". These do not seem like a particularly complicated or implausible set of claims. If you want to claim that AI risk depends on a way longer list, you’ll need to say what’s on the list. 11. “Anthropomorphic analogies probably mislead” Conceded.
Show more
the DoW attacking EAs is clearly setting us on the path to the next great political realignment:
You can do genuinely fun investigations of the AI risk web because they really are a sprawling ecosystem of tens of thousands of complexly interconnected scientists, policy experts, educators, etc. If you do this for the other side it's lame, "oh it's all just a16z employees"
Show more
Claude and ChatGPT trying to pin down my vibe based on interrogating me and surveying people who know me vs. Grok trying to pin down my vibe based on my tweets I can't wait for this international ASI ban to happen so I can tweet about other fucking things
Show more