If you suddenly find yourself concerned about AI and would like to transition to working in AI safety from a different field but aren't sure where to start, please feel free to reach out!
I have a tiny following on this website so not many are likely to see this - but I encourage my colleagues working in AI safety to say similar things.
I expect there will be a lot of people with useful skills who will be looking to get involved after the events of the last couple of weeks.
Show more
I’m very concerned that during RSI, labs will just stop externally deploying their models.
Which means they'll be going full steam ahead on the most dangerous use case of these models (recursive self-improvement), while the public remains in the dark about the nature of capabilities and the state of alignment.
And we end up on a path towards tremendous concentration of power.
Show more
Don't worry guys, WSJ says these three guys + Claude and Codex only got the "secret sauce" (OpenAI's algorithmic secrets) not the "crown jewels" (the weights)!!
I highly doubt Chinese threat actors haven't gotten further in. Racing faster might only give us the illusion of a lead! If we were serious about securing the frontier we'd follow through on export controls + pace so labs make wiser trade-offs on speed vs. security
Show more
On July 25, we hacked OpenAI.
Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc.
We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
Show more
Literally a year’s worth of changing attitudes about AI in one week
As I've said, I think there's probably some truth to the idea that AI companies may be incentivized to inflate the dangers of their product for financial and strategic reasons. However, when all of the major players in an industry are telling us, flat out, that the thing they're building could be very dangerous, it seems ridiculous to write their warnings off completely. If a guy built a weird contraption, handed it to me, and said, "Careful, it might kill you," I'm going to tend to take his warning seriously. It might be a hoax, sure, but if I know considerably less about the contraption than the guy who built it, I'm probably not going to just assume it's a hoax.
I'm not aware of any example in history where almost everyone INSIDE a particular industry is united in saying that their own product might be catastrophic for mankind. The climate alarmists have said that about the fossil fuel industry, for example, but the climate alarmists are not inside the industry. Notably, the fossil fuel industry itself has generally disagreed with their assessments. That's why the climate alarmism comparisons here just don't hold water. It's a totally different situation.
So, while we should be cautious about taking anything at face value, I find the firm "there's nothing to worry about" claims to be pretty ridiculous. This is an extremely powerful and rapidly evolving technology. Any powerful technology carries risk. If the people building this technology say the risk is high, again it seems foolish to blithely assume that all of them are conspiring to simply lie about the dangers of their own product.
Another read on the situation -- one that allows you to doubt the benevolence of these people (as I most certainly do), while not necessarily discounting all of the hazards of the thing they're building -- is that they pushed full speed ahead on this technology, blew right past the warning signs, and now they're seeing things on their end that legitimately scare them, and the PR blitz in the media right now is a classic case of corporate CYA. When the shit hits the fan, they can say, "We tried to stop it." Even though they actually could have stopped it, but didn't, when they had the chance. That seems like a plausible scenario, too.
Show more
I'm usually a pretty private person, but given my face has been in the news(!) I wanted to share a bit about myself.
Before moving to SF I was a British diplomat. I worked on missile defense and CBRN defense at
@UKDELNATO during the Russian invasion of Ukraine. I moved to SF to work more directly on making sure humanity gets to maximally benefit from the paradigm-shifting technology that is AI.
I worked at Open Philanthropy (who were early to AI Safety) and FutureHouse (who are trying to build “The AI Scientist”). I have always wanted the world to benefit from the full potential of this technology.
I want to make sure that we navigate the next few years and decades as safely as possible. I want us to maximize economic prosperity and for democratic values to flourish. And that is why I feel such pride about the work we’re doing at METR.
There are strong business incentives for the developers of frontier AI to not be maximally transparent, and whilst I have no ideological commitment to who plays this role, there need to be experts outside of companies providing the public some transparency into what they are building.
METR’s culture isn’t intellectually homogenous. Disagreements are encouraged. I work next to critics and skeptics of regulation as well as advocates. To spend a day in the METR office is to understand that, above all else, we care about Doing Science Well. How does one understand model capabilities? How can we make sure these models are controllable and monitorable?
So yeah, that’s a bit about me! Here’s a way better photo than the one they splashed on the front page of the Post. Scandalously, it was taken in London.
Show more
👀 SCOOP: U.S. open to discussing AI "shared risks" with China, Bessent says
Americans deserve representation in Congress that is not only well-informed on cutting-edge AI development, but also has the wherewithal to cross the aisle and overcome political pressures to address the concerns these technologies pose. We need public hearings now.
Show more
a lot of people work at METR so you can do the left version of this where they're all white men too if you want:
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why.
Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not.
We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up).
METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website.
Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Show more
They left me out because a white dude with glasses had too much risk of Watney
0. The core disagreement was about the inevitability of a race
1. I think leadership is way too paranoid about China and the US government. They don’t believe it will be possible to negotiate.
2. They largely initiated the recent race to RSI, because of a belief in its inevitability. Note that OpenAI had to shed a bunch of dead weight like Sora because Anthropic was going for the jugular.
3. Even if they are **not** being pessimistic, I disagree with their consequentialist philosophy. If the race is inevitable you should not contribute.
Show more
There’s an impulse in American politics, a set of tactics and drives that has proven very effective at extinguishing speech. We once called it McCarthyism. Another time we called it Wokeism. One day in the future we will have some name for the campaign that is underway against all people who believe in serious AI risks, and whatever that name ends up being, it will be yet another name for that old impulse.
It’s an ugly one. I recommend you avoid being a part of it, no matter your politics or your thoughts on AI. I suggest you think for yourself. You really do want to worry about the state projecting too much power over AI, centralizing control or robbing us of the unambiguous benefits of the technology in the name of preserving the economic status quo. But there really are grave risks from AI that go beyond what any other mass-market digital technology has posed.
These two things aren’t totalizing worldviews that exist in conflict. You don’t need to believe only one or the other. *Both of them are true.* The question should not be “which facts should we ignore, and which should we pay attention to?” The question is: “how do we walk the narrow corridor between all these realities which are in deep tension with one another?”
You should think for yourself. Assess other people’s work for yourself rather than allowing powerful people to put a label on them for you. As someone who has been a journey that has taken me to different perceived “sides” of the AI debate, believe me: You’ll find that there are bright and thoughtful people on both the “safety” and “accelerate” sides, and you will also find that both sides have their imbeciles, charlatans, and, occasionally, true cretins. It’s your job to figure out what’s what. Don’t let other people think for you. Don’t be played a fool.
Show more
I’ve been warning about this since 2000 ⤵️
"For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability"
Seems like a reasonable idea. Any reason why these can't be public? Or at least give third parties access who could then publicly discuss to what extent they find the safety cases persuasive?
Show more
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot.
We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors).
Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process.
Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases.
We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these.
When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs.
Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring.
Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
Show more
When we met Jason, we took to calling him a "happy warrior". He's indescribably wise, battle-tested from 20 years serving in the Air Force, never short on tales about negotiating with the Saudis and being on the inside of DTRA - with a heart of gold and a deep sense of duty to his sons that now compels him to work on AI safety. Jason is going to do wonders for our federal/natsec engagement. Congrats to us!
Show more
After 20 years in the Air Force, and a summer as a fellow at
@GovAIOrg , I’m excited to get to work as the Director of NatSec Policy at
@EncodeAction alongside the amazing team led by
@SnehaRevanur!
…and hoping to become half as prolific on here as
@_NathanCalvin
Show more
Encode does super important work in AI policy. Check this opportunity out!