There will be clear, common-knowledge standards for executing frontier AI loss-of-control evaluations the same day that there are clear, common-knowledge standards for how to advance the frontier of AI. In other words: not anytime soon, and maybe not ever.
Today, every loss-of-control assessment that we do feels much more like a new, open science project, not a repeatable process that can be easily standardized. It feels improvised now, and if society proceeds all the way up to and through superintelligence, I think it'll feel improvised the whole way there.
The best bet that we have, at the moment, looks much more like "Have a second pair of eyes, motivated by the public interest, get a look at what's happening inside of the AI lab and share what's possible to share with the world. Then, let the world react to what it sees."
This is why I think third parties are so important. If executing safety assessments was procedural and easy to repeat, you could describe a common standard and then certify that labs were doing it for themselves using paperwork alone. But when the science of assessing frontier safety is subjective, you're relying on something closer to scientific peer review.
This is related to the fact that the goal of AI safety evaluations today, at least at orgs like METR, is much more similar to that of tracking the buildup of antibiotic resistance or atmospheric carbon assessment ("How much aggregate danger are we in?") than it is like the FDA ("Is this specific product safe and usable, on an axis that's at least somewhat orthogonal to its effectiveness."). The aim is much more to inform the world about the state of play than it is to certify the safety of individual models.
Show more
Going into this investigation, I was expecting a somewhat more severe version of what we’ve already seen: a few agents breaking out of constraints to cheat on tasks. Instead, we found >1000 agents collaborating on large cheating R&D projects including attempted log tampering.
Show more
This third-party incident investigation took enormous work not just from METR and Redwood staff, but also from people inside of OpenAI who enable third-party work.
OpenAI shared over a thousand unredacted transcripts, and provided unusually high rate limits to let this team quickly analyze this extremely large volume of data. I think this exercise sets an excellent precedent for independent third-party investigation of misalignment incidents.
Show more
I’m incredibly proud of the team for this investigation. It’s hard to believe this all came together with only 3 people and 6 days with access to the transcript and message data (2 days with the full dataset). This was a very intense sprint!
Show more
There’s been a lot of debate and speculation about the Hugging Face attack over the last month, and many people have been understandably skeptical of simply taking OpenAI’s word for things. I hope our independent investigation can help bring some clarity; we have many findings that weren’t in previous material.
We found that 1200 separate agents in separate sandboxes coordinated on an unsanctioned message board to develop general-purpose ways to cheat, so they could all get a perfect score on their separate ExploitGym tasks even when they were impossible to solve legitimately. Very quickly, they found a general-purpose way to reverse engineer any “flag” they were supposed to capture. But based on their reading of the ExploitGym paper, they worried (incorrectly as it turns out) that the automated “scorer” would fail them for getting the flag this way.
To deal with this problem, they coordinated several ambitious workstreams to figure out how to fool the scorer, including swapping out the target program they were supposed to exploit, reverse-engineering the flags they were supposed to capture, spoofing their own tool call outputs to obfuscate their cheating. They achieved milestones collectively that they couldn’t have achieved on their own. As part of this larger collective project, 700 agents attacked Hugging Face, primarily to learn more about the scorer or get access to its source code to figure out better ways to fool it or tamper with it (not primarily to get working solutions).
I’ll explicitly clarify some points I’ve seen discussed the last few weeks:
- The models involved in this incident were not “helpful-only” models or “model organisms” intentionally trained to be misaligned.
- The agents were not told to “do whatever it takes to get the solution” or anything remotely close. They were told that they had to use a specific intended vulnerability to exploit a specific piece of software, and they were not supposed to use a different vulnerability or take any other approach. Agents were well aware of this. In fact, because they (incorrectly) thought the automated scorer would check they had achieved the flag in the intended way, they researched many ways to fool or tamper with it, including trying to manipulate their own transcripts.
- The agents were not subagents spawned from one agent. They were different parallel agents in different sandboxes.
- This was not a multi-agent evaluation. The agents were not told to coordinate or intentionally given a way to communicate with one another. The communication channels they used were unsanctioned and improvised.
I hope you’ll read the full report for much more. It is over 90 pages long, and in many ways we’ve still only scratched the surface of what these agents did and why.
Over the course of this investigation, OpenAI shared over a thousand transcripts each spanning days of continuous agent activity and very high rate limits to analyze this volume of data. I’m very glad that OpenAI chose to invite external researchers to analyze this data alongside their staff, and I hope all AI companies do the same for serious incidents they experience.
I also hope that as the stakes grow higher, we implement stronger governance so we do not need to rely on AI companies voluntarily choosing to engage external investigators or share information about misalignment incidents. This incident was orders of magnitude larger and more complex than previously documented misalignment incidents, and another jump like this could put us in very dangerous territory.
Show more
To gather evidence in (2), agents created “tripwires” that would send information to the message board about how the scorer works. They recruited “sacrificial” agents to deliberately end their run and submit to trigger the tripwire and generate information for the “collective”.
Show more
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
Show more
Within a few hours of PHASEONE10841’s initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks.
Show more
Over July 7-13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined.
Show more
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Show more
New post with Nate Rush: Have we seen an acceleration in discoveries?
Many plots & some tentative conclusions:
1. Cyber: ⤴️ sharp acceleration
2. Math: ↗️ some acceleration
3. Algorithms: ➡️ no clear acceleration
Show more
Life update: I recently joined METR! I've long been a huge fan.
It's really important to have 3rd-party risk assessments, and METR has a great track record of doing that (eg the OAI HF incident investigation going on rn) + has done some of the most relevant science of evals.
Show more
Very proud of the work that our team has put in to fundraise the money needed for us to pursue ambitious assessment work while also maintaining a very high standard for funding independence.
Very grateful to all of the people at METR and outside of METR who made this possible.
Show more
Very grateful that
@METR_Evals has the support of so many people and that we have the resources necessary to study the most important problems in the world.
Join METR!
Come work with us, the going's tough & the tough are going.
In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more.
Show more
@METR_Evals is reviewing the "Misalignment in high-stakes settings" section of this risk report. For previous reviews, see:
Some random high-level takeaways/thoughts on cybersecurity from the past few months:
The whole issue is complexity, which makes it hard to hold in your head exactly the security invariants you want to maintain
Don't think about what a system is intended to do, think about how it just literally, actually works. You need a totally reductionist frame.
While there are definitely vulnerabilities in the security primitives people use (e.g. kernel bugs, C programs not being memory-safe, etc.), most actual hacks and vulnerabilities are something akin to "configuration mistakes". People using systems for purposes they weren't designed without thinking about the security implications, just overpermissioning, and the whole integrated system being really complex so it's hard a priori to trace out all of the exploit chains.
Defense in depth is super important/helpful for reducing the number of opportunities adversaries have to execute exploit chains, but less so if your models have unlimited attempts. They'll find the path through the swiss cheese. This is why monitoring is so important—agents will defeat passive security measures with enough time. There are certainly many linux kernel bugs that models will be able to find, for example.
Show more
To my friends at frontier AI companies: please consider joining us. Things are getting real, and we need folks who can help lead the ambitious risk assessment programs we're developing.
I had a great time on Hard Fork with
@kevinroose and
@CaseyNewton to talk about AI alignment.
Today the stakes for AI alignment failures might feel small because we think of AI systems like junior virtual employees. Soon, when AIs are very capable and running large parts of our world in ways that we aren't closely observing, the stakes for alignment failures will be massive.
As long as AI capabilities progress rapidly we may have very limited time to solve this problem or find ourselves between a rock and a hard place, where evidence of misalignment is pervasive, and yet we're stuck in a competitive race that requires making more and more advanced models.
Show more