Before succumbing to the temptation to naval gaze into the political theory abyss, it's worth stepping back and clarifying what exactly is happening and being proposed.
Several US companies are on the precipice of fully automating the AI R&D loop, inclusive of pre/post training, env creation, data generation, evals, algorithm and kernel design, systems engineering, architecture search, etc. -- the full stack.
We are already in a regime of weak RSI via partially automated SWEs, but closing the loop altogether represents a difference in degree becoming a difference in kind. The pace of progress will be explosive and potentially uncontrollable.
The US companies closest to this threshold are warning that they are unprepared for a runaway intelligence explosion, and yet feel locked into a prisoners dilemma vis a vis each other and to a lesser extent vis a vis China.
We've already seen how rapid and comparatively unbounded progress is in verifiable RL domains, leading to spikey forms of superintelligence in math and cyber, including models that can prove open math conjectures, discover massive speed-ups for breaking encryption, and execute sophisticated multi-step exploits. We've also recently seen several severe examples of "loss of control" / misalignment incidents given inadequate monitoring and sandboxing practices relative to model capability. Moreover, these new capabilities mostly stem from scaling-up long-horizon post-training on legacy clusters, with OOMs of new compute about come online / in construction.
In the pre-RSI regime, human frictions created automatic buffers between new model releases, giving researchers and society time to probe emergent capabilities, design better evals, develop novel alignment techniques, and adapt / harden their infrastructure. As progress has accelerated, capability improvements have already started outstripping our adaptive capacity, as manifest in METR's inability to evaluate model autonomy beyond 13 hours, and narrow window for cyber defenders to prepare for open weight versions of Mythos.
RSI will exacerbate all these issues and create all new ones. At minimum, we should anticipate
- the equivalent of a GPT-5.2 -> 5.6 leap in capabilities at least every 24 hours (down from 3-6 months),
- concurrent algorithmic improvements densifying models to ultra-efficient sizes at any given capability level
- 100x Mythos-like capabilities across most verifiable domains, including chem, nuclear and bio
- new forms of multi-agent misalignment risk
- "company in a box" agents trained to stand-up whole organizations / corporations
- "cyber nuke"-like capabilities that require de minimis infra
- several transformer-scale breakthroughs, such as for long-term memory / continual learning, open-ended domains, and/or all-new training techniques for idealized "GPT-zero"-esque metalearners
- concurrent speedups in any complementary technical domain, i.e. explosive rates of R&D and novel discoveries
It seems to me there is little to lose, and much to gain, from having the social technology to "pace" these developments rather than to let them rip with zero industry / gov't coordination, particularly as there is technically no law explicitly prohibiting a company from letting an RSI loop run indefinitely and unleashing whatever comes out the other end into the world.
There are innumerable ways an uncoordinated intelligence explosion could become an unmitigated disaster for the cause of liberalism, including runaway power concentration, rapid societal destabilization, rogue AIs / loss of control scenarios, WMD mass proliferation, vulnerable world technologies, and beyond.
Human civilization is about to be forever changed regardless, however if were possible to coordinate the handful of key actors and create artificial "buffers" between each step-change in model capability to enable adaptation, mitigation and alignment research to catch-up, it's worth a shot.
Given the short-timeline, I think a DPA 708-style agreement is probably our best bet, i.e. an industry consortia with narrow antitrust carveouts for sharing safety and security practices, funding an assurance nonprofit / independent verification organization for 3rd party evals, incident reporting, internal deployment monitoring, standards setting, and enforcing a protocol for coordinated delays / slowdowns, among other things. This still leaves open the China question but that's a bridge we won't cross until after solving the collective action problem at home.
I'm open to other approaches / coordination frameworks but this is the object level issue we're facing. Political theory is great, and I would love to use our limited steering capacity to guide AI development toward a future that maximizes individual liberty, but as a discussion baseline, gesturing at philosophical abstractions is simply non-responsive to the crisis at hand. A red-herring at best, a suicidal circlejerk at worst.
Show more