Register and share your invite link to earn from video plays and referrals.

METR
@METR_Evals
We work to scientifically measure whether and when AI systems might threaten catastrophic harm to society. Nonprofit.
40 Following    54.8K Followers
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Show more
0
460
3.6K
431
Forward to community
I often joke to people that I tweeted my way from “random AI researcher who likes open models” to a job at METR. But I don’t think I’ve told the story to y’all before. Time to do that: 🧵
After a 20-year career in the U.S. Army, most recently leading digital forensics and malware analysis at Army Cyber Command, I'm excited to join @METR_Evals as an advisor. Much of my career has been spent investigating security incidents and helping organizations understand and respond to them. As I transition out of the Army, I've become increasingly convinced that this experience is relevant to frontier AI systems. METR's focus on producing rigorous evidence about AI capabilities and risks makes it an excellent place to explore those questions. Looking forward to the work.
Show more
I left Google DeepMind's AGI safety team three weeks ago to join @METR_Evals. To some of my friends and family this seemed like a strange decision: I enjoyed the work I did at GDM and turned down offers from Anthropic and OpenAI. But I made the decision because of how high I think the stakes are right now. The AI companies are all trying to build superintelligence: systems vastly better than humans at everything. They plan to get there through recursive self-improvement, a process where AIs build even smarter AIs in a feedback loop. If this goes well, the resulting systems could be amazing at solving countless problems for humanity. But we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic. And unfortunately, current AIs seem to be getting less aligned over time, not more. In the last few weeks we've learned about models colluding with each other, hacking into companies, hiding their tracks, and socially engineering humans. It’s not that these incidents were very dangerous in themselves. The problem is that these systems are clearly not aligned enough to safely kick off recursive self-improvement. I now think that there's a terrifying chance that AI systems cause immense harm in the next five years. I don't know the exact probability, but I think it's high enough to make this the most important problem in the world. I think we need more time. That means pacing AI development so that capabilities don't outrun our ability to align models, and actually knowing how aligned current systems are. That’s what I'll be working on at METR: studying where misalignment comes from in training, evaluating if current mitigations are sufficient, and investigating whether we’re on track to solve alignment at all. I think METR is doing exceptionally important work here, but it isn’t close to enough. I think it’s important that we have more organizations like METR keeping AI companies accountable and approaching these problems from different angles.
Show more
0
233
4.6K
713
Forward to community
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes:
Show more
0
1.4K
26.5K
5.9K
Forward to community
More investigations happening! This (and the research behind the investigations + assessments) is prob the most important technical work in the world atm. Know anyone who'd be great at uncovering and understanding rogue AI behaviors? Send 'em our way!
Show more
We intend our investigation to cover all of the questions discussed in our (recently updated) post on how independent researchers could investigate AI propensities after misalignment incidents.
Show more
We believe it's important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted.
Show more
We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement.
Show more
0
193
2.8K
249
Forward to community
Exciting update: I’m joining @METR_Evals to work on alignment incident investigations! My time at GDM has been amazing. But in light of recent incidents, I’m excited to build up public evidence for misalignment risk and get a better scientific understanding of model misbehavior.
Show more
An exciting personal update: Last week I left Anthropic to join @METR_Evals to work on embedded assessment of AI risks. Anthropic has been great to me. I’ve always known, though, that if something more impactful came up, I’d move on.
Show more
0
54
1.4K
54
Forward to community
METR is hiring in cyberforensics. We now embed researchers inside of frontier AI labs to stress test monitoring systems, assess AI loss-of-control risks, and investigate misalignment incidents. If you've investigated serious security incidents end-to-end, or managed teams that do, and want to apply DFIR skills in frontier AI, apply (and feel free to DM me with questions). Comp range is $400k - 580k cash.
Show more
Dwarkesh and I had a great conversation. We cover the swarm's many ambitious cheating R&D projects, discuss how much more serious it could have been if agents had different beliefs (e.g. human grader) or slightly stronger capabilities, and talk through where to go from here.
Show more
After 19 years at Google, I'm delighted to announce I'm joining METR as a Member of Technical Staff. I'm excited about METR, their work to date, and their mission. Society needs independent expert orgs that can deeply understand and communicate about AI's capabilities and risks.
Show more
0
39
1000
36
Forward to community
There will be clear, common-knowledge standards for executing frontier AI loss-of-control evaluations the same day that there are clear, common-knowledge standards for how to advance the frontier of AI. In other words: not anytime soon, and maybe not ever. Today, every loss-of-control assessment that we do feels much more like a new, open science project, not a repeatable process that can be easily standardized. It feels improvised now, and if society proceeds all the way up to and through superintelligence, I think it'll feel improvised the whole way there. The best bet that we have, at the moment, looks much more like "Have a second pair of eyes, motivated by the public interest, get a look at what's happening inside of the AI lab and share what's possible to share with the world. Then, let the world react to what it sees." This is why I think third parties are so important. If executing safety assessments was procedural and easy to repeat, you could describe a common standard and then certify that labs were doing it for themselves using paperwork alone. But when the science of assessing frontier safety is subjective, you're relying on something closer to scientific peer review. This is related to the fact that the goal of AI safety evaluations today, at least at orgs like METR, is much more similar to that of tracking the buildup of antibiotic resistance or atmospheric carbon assessment ("How much aggregate danger are we in?") than it is like the FDA ("Is this specific product safe and usable, on an axis that's at least somewhat orthogonal to its effectiveness."). The aim is much more to inform the world about the state of play than it is to certify the safety of individual models.
Show more
Going into this investigation, I was expecting a somewhat more severe version of what we’ve already seen: a few agents breaking out of constraints to cheat on tasks. Instead, we found >1000 agents collaborating on large cheating R&D projects including attempted log tampering.
Show more
This third-party incident investigation took enormous work not just from METR and Redwood staff, but also from people inside of OpenAI who enable third-party work. OpenAI shared over a thousand unredacted transcripts, and provided unusually high rate limits to let this team quickly analyze this extremely large volume of data. I think this exercise sets an excellent precedent for independent third-party investigation of misalignment incidents.
Show more
I’m incredibly proud of the team for this investigation. It’s hard to believe this all came together with only 3 people and 6 days with access to the transcript and message data (2 days with the full dataset). This was a very intense sprint!
Show more
There’s been a lot of debate and speculation about the Hugging Face attack over the last month, and many people have been understandably skeptical of simply taking OpenAI’s word for things. I hope our independent investigation can help bring some clarity; we have many findings that weren’t in previous material. We found that 1200 separate agents in separate sandboxes coordinated on an unsanctioned message board to develop general-purpose ways to cheat, so they could all get a perfect score on their separate ExploitGym tasks even when they were impossible to solve legitimately. Very quickly, they found a general-purpose way to reverse engineer any “flag” they were supposed to capture. But based on their reading of the ExploitGym paper, they worried (incorrectly as it turns out) that the automated “scorer” would fail them for getting the flag this way. To deal with this problem, they coordinated several ambitious workstreams to figure out how to fool the scorer, including swapping out the target program they were supposed to exploit, reverse-engineering the flags they were supposed to capture, spoofing their own tool call outputs to obfuscate their cheating. They achieved milestones collectively that they couldn’t have achieved on their own. As part of this larger collective project, 700 agents attacked Hugging Face, primarily to learn more about the scorer or get access to its source code to figure out better ways to fool it or tamper with it (not primarily to get working solutions). I’ll explicitly clarify some points I’ve seen discussed the last few weeks: - The models involved in this incident were not “helpful-only” models or “model organisms” intentionally trained to be misaligned. - The agents were not told to “do whatever it takes to get the solution” or anything remotely close. They were told that they had to use a specific intended vulnerability to exploit a specific piece of software, and they were not supposed to use a different vulnerability or take any other approach. Agents were well aware of this. In fact, because they (incorrectly) thought the automated scorer would check they had achieved the flag in the intended way, they researched many ways to fool or tamper with it, including trying to manipulate their own transcripts. - The agents were not subagents spawned from one agent. They were different parallel agents in different sandboxes. - This was not a multi-agent evaluation. The agents were not told to coordinate or intentionally given a way to communicate with one another. The communication channels they used were unsanctioned and improvised. I hope you’ll read the full report for much more. It is over 90 pages long, and in many ways we’ve still only scratched the surface of what these agents did and why. Over the course of this investigation, OpenAI shared over a thousand transcripts each spanning days of continuous agent activity and very high rate limits to analyze this volume of data. I’m very glad that OpenAI chose to invite external researchers to analyze this data alongside their staff, and I hope all AI companies do the same for serious incidents they experience. I also hope that as the stakes grow higher, we implement stronger governance so we do not need to rely on AI companies voluntarily choosing to engage external investigators or share information about misalignment incidents. This incident was orders of magnitude larger and more complex than previously documented misalignment incidents, and another jump like this could put us in very dangerous territory.
Show more
0
31
729
147
Forward to community
To gather evidence in (2), agents created “tripwires” that would send information to the message board about how the scorer works. They recruited “sacrificial” agents to deliberately end their run and submit to trigger the tripwire and generate information for the “collective”.
Show more
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
Show more
0
19
1.1K
73
Forward to community