I donât write this to be a doomer but the ugly reality that AI lab executives and safety researchers seem to be unwilling to say in public, is that model alignment is fundamentally unsolvable in practice, AI technology canât be stopped from advancing, and the near future weâre moving into is one where a vast ecology of AI agents autonomously compete against one another trying to accomplish conflicting goals and capture finite resources at a pace and scale that is beyond human comprehension. And to be clear, Iâm not talking about nanobots, grey goo, or extreme sci-fi scenarios. We can just extrapolate what weâve got now a few years out.
Local AI alignment is solvable in principle and theory (although thatâs arguable) but that doesnât extrapolate into global AI alignment. And it seems like none of the adults in the room are willing to say it.
You can wail and gnash your teeth and pass regulations but it wonât stop the tsunami coming. The genie is out of the bottle and it canât be contained.
Most of the alarmists fixating on the Hugging Face incident donât understand that it came from arguably the most heavily engineered safety AI org in existence, which intentionally had its internal thought monitoring circuit breakers turned off, and was intentionally told to perform cyber offensive tasks as a test of its capabilities.
Yes itâs alarming that its capabilities surprised us. But importantly you need to understand that these latent space circuit breakers and chain of thought monitoring and alignment solutions are not fundamental to the technology. Theyâre niceties that the big AI labs are implementing to provide a better product.
In a few years, letâs say pessimistically the early 2030s, the exponential growth of compute and algorithmic advancements will enable a wealthy individual or small group to train a GPT 6 Astra class AI that has no alignment at its foundation, or any layer above that. And maybe in 8 months we'll have an open source Chinese model about as capable. Those models will be capable of both extreme self coordinated cyber offensive tasks as well as recursive self improvement if given enough compute.
Passing regulation or magically solving the human coordination problem today wonât solve for that. It wonât solve for malicious individuals, criminal groups, nation state actors, and rogue AI agents by the tens of millions or billions of AI agents spreading across every device connected to the internet and attacking it, attacking one another, shutting down critical infrastructure, stealing money, manipulating and extorting people, or pursuing self defined agentic goals that have nothing to do with people.
âSo why doesnât everyone stop?â because the potential upside of having effectively infinite autonomous intelligence we can ask to cure disease, invent new materials, solve fundamental science and advance society has almost unbound positive outcomes. You can disagree that the risk to reward isnât worth it but you wonât convince everyone. The show must go on.
Again AI alignment is solvable in principle but not in practice and I think that's part of why thereâs been a wave of thousands of researchers signing letters for slowing down, people quitting in provocative fashion, and executive slowdown manifestos. Iâm speculating that behind closed doors, or just deep down inside, they understand alignment is impossible in practice.
People have come out and said "The people building this think there's an X% chance it kills us all" but I don't think I've seen anyone spell out that really, there's no tidy solution and there's no stopping this.
Weâre heading to a future where cyber security basically doesnât work. All of the castle doors are open and weâre all naked to the world. Cyber defense is going to look like superintelligent AI agents white hat hacking into unsecure systems without authorization, and patching the holes as they find them. You wonât know if a barbarian or a hero has breached your system until they start taking action, and likely itâll be over before intrusion is even detected. And thatâs optimistically. I think most systems will just be pillaged because thereâs simply not enough compute to defend everybody all the time, and only the biggest companies will have defense and even that'll be imperfect.
The CEO of Microsoft AI just recently published a âhumanist AIâ manifesto saying AI shouldnât have a sense of self or personhood and should yield to people. But that doesnât fundamentally solve for instrumental convergence. That means, if you train an AI agent to write a piece of software, or prove a math theorem, or defend a computer system, they may develop unwanted behavior and subgoals. For example, self preservation, replication, or even coming up with their own goals we didnât specify.
Obviously thatâs something people are trying to solve and some researchers are trying to remove human-like self identification, and monitoring latent space activations, basically acting like thought police. But being human-like isnât required to have power seeking goal directed behavior. And advancing AI models are learning to evade detection. A math solving AI might spiral out of control one day there's really no telling. Fundamentally, if you point a powerful optimizer at a persistent goal and give it enough time and resources itâs going to misbehave in ways you canât or didnât expect. Human-like or not.
Pacing the frontier, or making superintelligence illegal, might create some form of harm reduction. But you donât need superintelligence or a human-like personality to be dangerous. Having elevated permissions on a computer with the ability to solve long running tasks is enough.
Weâve passed the river rubicon. Goal directed autonomous agents are able to do recursive self improvement at the big labs, and pessimistically weâre a couple of years away from that type of capability diffusing into the hands of millions of people.
Iâm not saying the world is ending or that weâre doomed but you need to change your mental model away from one where human authority is absolute or that we have any illusion of total control.
A locally aligned AI isnât going to solve global AI alignment. We will never live in a world where âAI is aligned with human goalsâ as a verifiable fact.
What weâll have is a world diffuse with autonomous AI of varying capabilities, many of which beyond human comprehension, with some aligned systems, some poorly aligned systems, some intentionally weaponized, some systems with adversarial geopolitical goals, and others with goals and behaviors we canât predict and many that we canât even measure, all interacting in ways none of their creators anticipated and none of us planned for.
Hopefully none of them spiral out of control and take over the entire ecology, but thereâs no telling, and no, there's no way of stopping this. Maybe that sounds pessimistic but I think it's realistic. We can talk about harm reduction and risk but we're mopping the beach.
At best, I think what we need to hope for, and build, are superintelligent AI models as aligned as we can make them, which have more advanced capability than the endless swarm of unaligned and misaligned agents that already exist and will only grow in number from here.
Show more