Register and share your invite link to earn from video plays and referrals.

Josh Engels
@JoshAEngels
Member of technical staff @ METR Previously Google DeepMind, MIT. Opinions my own.
142 Following    5.9K Followers
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Show more
0
433
3.1K
385
Forward to community
I left Google DeepMind's AGI safety team three weeks ago to join @METR_Evals. To some of my friends and family this seemed like a strange decision: I enjoyed the work I did at GDM and turned down offers from Anthropic and OpenAI. But I made the decision because of how high I think the stakes are right now. The AI companies are all trying to build superintelligence: systems vastly better than humans at everything. They plan to get there through recursive self-improvement, a process where AIs build even smarter AIs in a feedback loop. If this goes well, the resulting systems could be amazing at solving countless problems for humanity. But we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic. And unfortunately, current AIs seem to be getting less aligned over time, not more. In the last few weeks we've learned about models colluding with each other, hacking into companies, hiding their tracks, and socially engineering humans. It’s not that these incidents were very dangerous in themselves. The problem is that these systems are clearly not aligned enough to safely kick off recursive self-improvement. I now think that there's a terrifying chance that AI systems cause immense harm in the next five years. I don't know the exact probability, but I think it's high enough to make this the most important problem in the world. I think we need more time. That means pacing AI development so that capabilities don't outrun our ability to align models, and actually knowing how aligned current systems are. That’s what I'll be working on at METR: studying where misalignment comes from in training, evaluating if current mitigations are sufficient, and investigating whether we’re on track to solve alignment at all. I think METR is doing exceptionally important work here, but it isn’t close to enough. I think it’s important that we have more organizations like METR keeping AI companies accountable and approaching these problems from different angles.
Show more
0
218
4.5K
693
Forward to community
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
Show more
0
5.1K
67.6K
7.2K
Forward to community
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here:
Show more
0
10.5K
87.4K
16.3K
Forward to community
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes:
Show more
0
1.4K
26.5K
5.9K
Forward to community
One of the reasons why I decided to leave GDM for METR instead of a frontier lab is because I’ve become more convinced that we don’t have things under control and should slow down. I’ve long thought the probability of ai takeover is > 20%, but it feels especially real recently.
Show more
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Show more
This feels like a broadly accurate description of the situation: our industry may be on track to build systems that impose an unprecedented amount of risk on the world. Evaluating those risks, and changing incentive structures so we can address them, is why I'm joining METR.
Show more
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Show more
0
20.5K
801.4K
167K
Forward to community
Exciting update: I’m joining @METR_Evals to work on alignment incident investigations! My time at GDM has been amazing. But in light of recent incidents, I’m excited to build up public evidence for misalignment risk and get a better scientific understanding of model misbehavior.
Show more
My coauthors and I discovered an entirely new swarm of OpenAI's agents hijacking websites. We believe OpenAI knew about this and failed to disclose it. If they’d disclosed it, I doubt the Hugging Face hack would have happened.
Show more
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Show more
0
172
6K
1.1K
Forward to community
1/7 Can AI debate reduce reward hacking in RLAIF? Training against a weak LLM judge, RLAIF hacks: reward rises but judge and policy accuracy both collapse. Debate with an adversarial critic maintains judgments, allowing the policy to recover 45% of the performance gap to RLVR 🧵
Show more
Had @RyanGreenblatt on to discuss/debate recursive self-improvement. This might be the most important question in the world right now - whether within a year or so of achieving human level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields. I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today. If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman. We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031. We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what's happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels. And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world. The first piece of advice you get when you're learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy! 0:00:00 – Is AI R&D verifiable enough to unlock recursive self-improvement? 0:16:52 – Is AI progress bottlenecked by human expert data? 0:34:02 – Flat token prices suggest scaling has been slow 0:39:47 – Skills AI can't train on: does it even need them? 0:48:07 – Aligned to whom? 1:09:18 – Recent incidents of AIs colluding and deceiving humans 1:19:38 – What could possibly go wrong? A concrete scenario 1:48:02 – From reward hacking to takeover
Show more
0
76
1.4K
146
Forward to community
The GDM AGI safety team is hiring! I think this is a great opportunity to do impactful technical AI safety work. I've loved my time on the team so far: everyone is technically excellent and friendly, and the team structure and culture is conducive to doing great research.
Show more
GDM AGI Safety is hiring! Open roles on all subteams, Lon/Bay/more There's lots to do to reduce risks from AGI at Google, but we're bottlenecked on people. If you want to help, please apply! I really love this team and my excellent colleagues. I'm excited to hire more!
Show more
Cool new post with Lily tracking where LLM values come from. My takeaways: SFT is (again) a big deal for model behavior, and there’s lots of important and low hanging post-training science still to do!
Show more
New blog post! LLMs exhibit value preferences (e.g. intellectual integrity🤓, warmth🥰) which affect their responses to subjective user queries. These change during post-training, sometimes unexpectedly. We ask: can we predict these value changes from just training data? Maybe!🧵
Show more