I left Google DeepMind's AGI safety team three weeks ago to join
@METR_Evals. To some of my friends and family this seemed like a strange decision: I enjoyed the work I did at GDM and turned down offers from Anthropic and OpenAI. But I made the decision because of how high I think the stakes are right now.
The AI companies are all trying to build superintelligence: systems vastly better than humans at everything. They plan to get there through recursive self-improvement, a process where AIs build even smarter AIs in a feedback loop. If this goes well, the resulting systems could be amazing at solving countless problems for humanity. But we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic.
And unfortunately, current AIs seem to be getting less aligned over time, not more. In the last few weeks we've learned about models colluding with each other, hacking into companies, hiding their tracks, and socially engineering humans. It’s not that these incidents were very dangerous in themselves. The problem is that these systems are clearly not aligned enough to safely kick off recursive self-improvement.
I now think that there's a terrifying chance that AI systems cause immense harm in the next five years. I don't know the exact probability, but I think it's high enough to make this the most important problem in the world.
I think we need more time. That means pacing AI development so that capabilities don't outrun our ability to align models, and actually knowing how aligned current systems are. That’s what I'll be working on at METR: studying where misalignment comes from in training, evaluating if current mitigations are sufficient, and investigating whether we’re on track to solve alignment at all.
I think METR is doing exceptionally important work here, but it isn’t close to enough. I think it’s important that we have more organizations like METR keeping AI companies accountable and approaching these problems from different angles.