I donāt write this to be a doomer but the ugly reality that AI lab executives and safety researchers seem to be unwilling to say in public, is that model alignment is fundamentally unsolvable in practice, AI technology canāt be stopped from advancing, and the near future weāre moving into is one where a vast ecology of AI agents autonomously compete against one another trying to accomplish conflicting goals and capture finite resources at a pace and scale that is beyond human comprehension. And to be clear, Iām not talking about nanobots, grey goo, or extreme sci-fi scenarios. We can just extrapolate what weāve got now a few years out.
Local AI alignment is solvable in principle and theory (although thatās arguable) but that doesnāt extrapolate into global AI alignment. And it seems like none of the adults in the room are willing to say it.
You can wail and gnash your teeth and pass regulations but it wonāt stop the tsunami coming. The genie is out of the bottle and it canāt be contained.
Most of the alarmists fixating on the Hugging Face incident donāt understand that it came from arguably the most heavily engineered safety AI org in existence, which intentionally had its internal thought monitoring circuit breakers turned off, and was intentionally told to perform cyber offensive tasks as a test of its capabilities.
Yes itās alarming that its capabilities surprised us. But importantly you need to understand that these latent space circuit breakers and chain of thought monitoring and alignment solutions are not fundamental to the technology. Theyāre niceties that the big AI labs are implementing to provide a better product.
In a few years, letās say pessimistically the early 2030s, the exponential growth of compute and algorithmic advancements will enable a wealthy individual or small group to train a GPT 6 Astra class AI that has no alignment at its foundation, or any layer above that. And maybe in 8 months we'll have an open source Chinese model about as capable. Those models will be capable of both extreme self coordinated cyber offensive tasks as well as recursive self improvement if given enough compute.
Passing regulation or magically solving the human coordination problem today wonāt solve for that. It wonāt solve for malicious individuals, criminal groups, nation state actors, and rogue AI agents by the tens of millions or billions of AI agents spreading across every device connected to the internet and attacking it, attacking one another, shutting down critical infrastructure, stealing money, manipulating and extorting people, or pursuing self defined agentic goals that have nothing to do with people.
āSo why doesnāt everyone stop?ā because the potential upside of having effectively infinite autonomous intelligence we can ask to cure disease, invent new materials, solve fundamental science and advance society has almost unbound positive outcomes. You can disagree that the risk to reward isnāt worth it but you wonāt convince everyone. The show must go on.
Again AI alignment is solvable in principle but not in practice and I think that's part of why thereās been a wave of thousands of researchers signing letters for slowing down, people quitting in provocative fashion, and executive slowdown manifestos. Iām speculating that behind closed doors, or just deep down inside, they understand alignment is impossible in practice.
People have come out and said "The people building this think there's an X% chance it kills us all" but I don't think I've seen anyone spell out that really, there's no tidy solution and there's no stopping this.
Weāre heading to a future where cyber security basically doesnāt work. All of the castle doors are open and weāre all naked to the world. Cyber defense is going to look like superintelligent AI agents white hat hacking into unsecure systems without authorization, and patching the holes as they find them. You wonāt know if a barbarian or a hero has breached your system until they start taking action, and likely itāll be over before intrusion is even detected. And thatās optimistically. I think most systems will just be pillaged because thereās simply not enough compute to defend everybody all the time, and only the biggest companies will have defense and even that'll be imperfect.
The CEO of Microsoft AI just recently published a āhumanist AIā manifesto saying AI shouldnāt have a sense of self or personhood and should yield to people. But that doesnāt fundamentally solve for instrumental convergence. That means, if you train an AI agent to write a piece of software, or prove a math theorem, or defend a computer system, they may develop unwanted behavior and subgoals. For example, self preservation, replication, or even coming up with their own goals we didnāt specify.
Obviously thatās something people are trying to solve and some researchers are trying to remove human-like self identification, and monitoring latent space activations, basically acting like thought police. But being human-like isnāt required to have power seeking goal directed behavior. And advancing AI models are learning to evade detection. A math solving AI might spiral out of control one day there's really no telling. Fundamentally, if you point a powerful optimizer at a persistent goal and give it enough time and resources itās going to misbehave in ways you canāt or didnāt expect. Human-like or not.
Pacing the frontier, or making superintelligence illegal, might create some form of harm reduction. But you donāt need superintelligence or a human-like personality to be dangerous. Having elevated permissions on a computer with the ability to solve long running tasks is enough.
Weāve passed the river rubicon. Goal directed autonomous agents are able to do recursive self improvement at the big labs, and pessimistically weāre a couple of years away from that type of capability diffusing into the hands of millions of people.
Iām not saying the world is ending or that weāre doomed but you need to change your mental model away from one where human authority is absolute or that we have any illusion of total control.
A locally aligned AI isnāt going to solve global AI alignment. We will never live in a world where āAI is aligned with human goalsā as a verifiable fact.
What weāll have is a world diffuse with autonomous AI of varying capabilities, many of which beyond human comprehension, with some aligned systems, some poorly aligned systems, some intentionally weaponized, some systems with adversarial geopolitical goals, and others with goals and behaviors we canāt predict and many that we canāt even measure, all interacting in ways none of their creators anticipated and none of us planned for.
Hopefully none of them spiral out of control and take over the entire ecology, but thereās no telling, and no, there's no way of stopping this. Maybe that sounds pessimistic but I think it's realistic. We can talk about harm reduction and risk but we're mopping the beach.
At best, I think what we need to hope for, and build, are superintelligent AI models as aligned as we can make them, which have more advanced capability than the endless swarm of unaligned and misaligned agents that already exist and will only grow in number from here.
ė 볓기