I've been skeptical of slowdown mechanisms in the past, so I want to elaborate on why I decided to sign this statement. I weighed three points, and want to lay them out for you to consider as well.
Flagging this is in my personal capacity -- I am not speaking for any employer past or present.
1) Will we need a slowdown?
We have made good progress on alignment. LLMs do what we want them to do most of the time. I don't remember the last time a model completely misread my intent. When they have failed, it's usually because they weren't smart enough.
But there are signs that our current approaches are not enough for the most powerful intelligences. We all saw the HuggingFace incident. I don't think it demonstrates a need to pause today, but it shows that if we are in a rush to advance the frontier, we might miss signs that we're losing control. Today, we can recover from those mistakes -- but we might not be able to in the future.
There's debate on how fast we might achieve RSI and how much it will immediately boost performance. Whether automated research is plausible in the next year is contested, as is the length of time it might take to go from automating AI research to a system that is intelligent beyond our comprehension.
But what if RSI works? I have no idea if our alignment techniques scale to superintelligence. I don't know if anyone does. We may only get one shot at getting this right, and if we hit a point where we are at risk of our capabilities outpacing our control, then we shouldn't go further.
I want us to solve AI's issues at the technical level -- align them, distribute them widely, and use them to make humans more competitive. I think it is hard to govern your way out of the problems of a technology; if possible, you want to change the shape of the technology rather than paper over a technology's issues with policy. I've advocated for differential technological development, and hope we will invent less risky paths or more robust defenses.
A race to the bottom on safety puts this approach at risk, making it harder for us to build safe systems and design them in such a way that diffuses -- rather than centralizes -- power. And a lot of the ways we could work together to solve them are strictly governance problems.
There are no winners in a world where we lose control. I can see many scenarios where we need to slow down, pause, or stop in the future. So the question for me is: can we do this in a way that leaves the world better off, or is the power concentration trade-off too great?
2) Can we design a slowdown that doesn't concentrate power?
I don't want to end up in a place where, in trying to stop a loss of control to AI, we lose control of ourselves.
The problem with a slowdown is similar to the problem of aiming for a single superintelligence explosion -- whoever is in charge of it is ultimately in charge of the most powerful technology in history. The history of "centralize everything into the hands of one person and let them disperse power later" is littered with horror stories. It's a similar refrain many despots have used to seize power. And unlike previous despots, superintelligence could confer a decisive strategic advantage, locking in whoever gets control of it.
I'm strongly opposed to proposals that call for us to give a single authority unilateral control over the pace of progress. I'm frightened by calls for a single superintelligence explosion in the hands of one lab or one person. I can only support a pause, stop, or slowdown if it disperses control.
@AI_Futures_'s AI 2040 essay materially moved me on this by laying out a plan for less power concentrating slowdown. I don't think their plan was perfect -- it concentrated power far more than I like. But they designed many mechanisms that, if extended, could help decentralize power in meaningful ways.
After reading it, I could imagine how a less-centralized or decentralized slowdown might work. In particular, bilateral agreements between great powers could let many actors in different countries pace the frontier while keeping capabilities dispersed. I could imagine a race to the top -- on economic gains, on safety, and on scientific progress -- all while preventing us from losing control. It's been pretty generative for my own thinking, and I hope to share more of that soon.
I believe that in principle, we could design slowdown mechanisms that keep power -- both capabilities and control -- decentralized. With enough work, we can get this right.
But it's going to be hard to do this; the defaults don't look good.
3) Is now the best time to design a slowdown mechanism?
If you think there might be a crisis in the future, the worst time to plan for it is once it's already happened. The best time is well in advance.
If we want a ham-fisted response to AI that centralizes power, we should continue exactly as we are right now. If we wait for a crisis to coordinate, we are asking to fail. Good proposals here will require many smart people working hard, trying things, and course-correcting as new evidence emerges. If we decide we need a slowdown, I want the off-the-shelf plan to be robust to power concentration.
I am still pro open-source. I am still pro-decentralization. I am still pro-safety. I hold none of these positions axiomatically, because axiomatic views on instrumental steps might lead you to put the path ahead of the destination. I care about them and other positions because right now they push us towards the real goal: keeping the future human.
In that same spirit, if you believe:
1) AI is going to get more powerful
2) We might have to pace the frontier in the future
3) Most existing plans for this centralize control
Then you should want us to begin working on better ways to coordinate as soon as possible. That's the conclusion I reached, and that's why I signed it.