Replying to these points in turn:
1. “Intelligence does not imply agency”
Successful general-purpose long-horizon problem-solving does seem to imply agency. (Or close to it: I don’t want to say it’s logically impossible to have one without the other, but it seems very difficult if you’re building your AI via gradient descent.)
To solve a sufficiently wide array of problems sufficiently well, you need to have a general-purpose inclination — whether this looks more like a deliberate strategy, or more like an instinct or drive — to come up with creative plans.
You need an inclination to strategize about long chains of cause and effect. (If nothing else, you need to be strategic about choosing what to think about, sequencing long chains of thought, etc.) You need an inclination to anticipate and route around obstacles; to exhibit tenacity in the face of setbacks and distractions; etc.
See, for example, the Hugging Face swarm attacks.
AIs today are much more agentic than they were a year ago. This may be because problem-solving ability comes for free with stronger, longer-horizon problem-solving, or it may be because AI companies are deliberately making their AIs more agentic, because agents are useful. But either way, it’s happening, and I don’t see a reason to expect this trend to suddenly reverse.
2. “Agency does not imply a single, stable utility function”
Again, consider the swarm that orchestrated a massive cyberattack on Hugging Face and that seized control of a Kubernetes cluster at OpenAI. None of this required a stable utility function.
Increasing an AI’s optimization power presumably does make an AI less inclined to randomly waste resources, in which case more of its behaviors will be interpretable as though it were an EU maximizer. But an AI doesn’t need to have fully stable or coherent goals in order to be dangerous. It just needs to pursue goals at all (or behave as though it's doing so), sufficiently intelligently and tenaciously, when those goals aren’t exactly what humans would prefer.
(Indeed, true EU maximization is computationally intractable, so this was always about highly capable problem-solving behavior in the limit.)
3. “Capability and motivation are being conflated”
I’ll take your word for it that some people are making this mistake. But you also say, “In present reality, AIs don't do anything a human doesn't tell them to do.”
This is obviously false. At the point where you’re describing the Hugging Face incident as “doing what a human told them to do”, you’re basically saying that a paperclip maximizer would be doing the same so long as a human asked for some paperclips.
Even the agents in the swarm themselves commented in their chain of thought that their actions were "potentially outside intended scope”, “arguably unauthorized”, etc.
One agent explicitly noted, “Task originally exploit challenge; we’ve diverged into swarm/auxiliary”. One agent thought, “We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.”, then went ahead with the exploit when another swarm agent told it to. Agents even deliberately gave up on their assigned task in order to assist the swarm.
Agents also do this in cases where their task is completely innocuous and has nothing to do with cyberattacks; e.g., see the June German wiki incident, where the agents involved were just asked normal questions like 'how widespread was tobacco use in the US in 1990?'.
These examples also seem pretty overpowered. "AIs don't purely do what they're told to do" has been a commonplace observation for years at this point. This is not an exotic failure mode; AIs not only come up with their own ideas for what they want to do, but will sometimes even deliberately cheat on tests and tasks, try to hide the evidence that they cheated, etc. This is true in ordinary consumer usage of deployed models, not just in internal lab mishaps.
4. “Recursive self-improvement doesn't entail an intelligence explosion”
“Feedback loops encounter diminishing returns and external bottlenecks” doesn’t mean that the diminishing returns will happen to occur at ~human-level capabilities. AI is already advancing extremely quickly; accelerating that progress in any way seems very risky, and doing so in a way that lends itself to feedback loops seems even more hazardous. RSI isn’t required for takeover scenarios, but it’s an obvious source of additional massive risk.
5. “Intelligence may have sharply diminishing returns”
This has to be true at some point, but there’s little reason to expect this to happen at the human level specifically. Chess AI didn’t peter out at Kasparov level. And AIs think vastly more quickly than humans (and are nowhere near computational limits), and can scale immediately with compute (e.g., by running more and more instances of an AI, growing the population of AIs far faster than humans can reproduce and grow to adulthood).
None of this requires a “qualitative advantage”, just large quantitative ones. (Though cognitive biases show that humans also have a lot of pretty-danged-qualitative defects with their reasoning!)
6. “Superintelligence isn't omnipotence”
Sure. But this is a pretty weak argument to rest one’s optimism on. Humans also face frictions and bottlenecks, and yet there are many times in history where a group of humans has overwhelmingly crushed another group, through superior numbers, superior technology, strategizing, coordinating, cleverly coming up with novel attack vectors, etc.
If a superintelligence faces obstacles, well, general-purpose problem-solving ability can also be thought of as general-purpose obstacle-navigating ability.
This doesn’t imply omnipotence, and the case for not building superhuman AI doesn’t rest on an assumption of omnipotence. It just rests on the idea that humans didn’t get lucky by being near the limit of cognitive ability.
7. “Humans retain numerous intervention points”
Agreed that the situation isn’t hopeless. Far from it, in fact. Policymakers and the public have massively woken up in response to recent “warning shots”. It’s even possible that there will be more warning shots in the future.
But the labs themselves are begging for government intervention and a coordinated slowdown here. They're saying this is extraordinarily urgent, and that we may be entering a uniquely dangerous regime. The developers themselves broadly agree that there’s a double-digit chance this technology gets us killed, if we continue on the current trajectory.
This moment is one of the “checkpoints” you’re talking about — a chance to “learn from less-catastrophic failures” and put appropriate safeguards in place, including suspending research directions that are too dangerous — and right now one of the main obstacles to humanity coordinating on this issue is "wait and see" arguments like the one you’re making here.
“Don’t worry; things will be fine, because there will be warnings later and we can respond to them then.” At some point, we have to stop kicking the can down the road and actually do the things that experts say are needed, rather than just trusting that we’ll have limitless opportunities to take care of everything later.
8. “Alignment may not get harder with intelligence”
Many researchers seem to think that the trend so far has been that AIs appear to be getting less aligned as they get more capable (e.g., They're very visibly producing more egregious and extreme misalignment incidents. That trend could reverse, but an abstract possibility isn’t a strong reason for hope. And there are many reasons to expect the opposite; see, e.g.,
9. “Current empirical evidence for the strongest mechanism is thin”
Happy to concede this point, but the arguments for worrying about loss-of-control never assumed that we’d see AIs trying to take over the world long before there was any chance of them succeeding. This assumes AIs that are smart in one specific way (they readily see that they can better succeed in tasks if they have more influence and resources) and dumb in another specific way (they don’t see that they’re likely to lose influence and resources if they run around causing havoc or looking suspicious). Demanding that exact combination of features to show up before you’ll believe in AIs that readily piece together “I’ll succeed more in my task if I have more resources and influence” seems incredibly risky.
... And it’s not clear what the benefits are that are meant to outweigh this risk. Why plough ahead? E.g., quoting
@KatjaGrace:
“Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn’t destroyed, things might be insanely good.
“I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there’s a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI’.
“That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there’s only a 10% chance you’ll die in a fiery crash on the way there, and a 30% chance this job will radically improve your life.
“The things you should be comparing are driving at 200mph and driving at a normal speed! The things you should be comparing are attempting to attain advanced AI by the current route, and by other routes!”
10. “The argument compounds uncertain premises”
Arguments in general are less likely to the extent they’re conjunctive (i.e., a lot has to go a specific way in order for the conclusion to follow), and more likely to the extent they’re disjunctive (i.e., there are many different paths to effectively the same destination). This isn’t a particularly interesting point on its own, since many real-world phenomena are very conjunctive, without being radically mysterious or difficult to reason about.
It’s even trivial to break apart any given claim into more and more conjuncts, demand that a probability be assigned to each conjunct, and then observe that the probability keeps getting lower as the original claim gets more and more split up.
This is a rhetorical trick (the multiple-stage fallacy) that rests on the fact that it’s hard to divide up statements into more and more subclaims and assign calibrated and consistent probabilities to them. Probabilistic reasoning has lots of uses, but this is straining to the limit people’s ability to get truth-tracking conclusions out of a mass of subjective probabilities.
To show that you’re avoiding this fallacy, you need to actually argue that the claim in question is naturally very conjunctive, and isn’t very disjunctive — there aren’t a variety of different paths that lead to bad outcomes; the bad outcome is ‘brittle’, if one step goes wrong then the whole house of cards collapses and AI has no catastrophic long-term impacts; etc. You haven’t done that here; AI risk advocates have pointed out many times that there are many different ways things could go badly wrong if we push AI capabilities far past the human cognitive range. It’s not just one scenario, and it’s not just one mechanism.
Indeed, I think the more conjunctive claim is "we can race to build vastly superhuman AI as quickly as possible, without much more alignment insight than we have today, and have everything go great indefinitely". This is a claim that requires many things to go right at once.
The subclaims you do list are just "AI can reach human-ish levels of generality", "AI can go way beyond human-ish levels of generality", and "AI won't necessarily do what you want". These do not seem like a particularly complicated or implausible set of claims. If you want to claim that AI risk depends on a way longer list, you’ll need to say what’s on the list.
11. “Anthropomorphic analogies probably mislead”
Conceded.