Register and share your invite link to earn from video plays and referrals.

Ajeya Cotra
@ajeya_cotra
Helping the world prepare for powerful AI. Risk assessment @METR_evals (opinions my own). Blogs: Planned Obsolescence (AI), Good Bones (whatever's on my mind).
502 Following    22.7K Followers
Dwarkesh and I had a great conversation. We cover the swarm's many ambitious cheating R&D projects, discuss how much more serious it could have been if agents had different beliefs (e.g. human grader) or slightly stronger capabilities, and talk through where to go from here.
Show more
It's funny that while we were recording, @RyanGreenblatt was in the middle of his 6 day sprint on the METR report, and already knew the counterexamples to all my objections about his takeover story, but obviously, he couldn't say anything lol. Would an AI really start some crazy conspiracy in order to pass an evaluation, where they try to build whole potemkin villages to fool the evaluator? And even if they did, why would other instances, who have different objectives, join the conspiracy? And even if they did, wouldn't at least some of the instances tattle on the conspiracy? It just seems crazy hard to sustain secret underground civilization inside an AI company, without humans and other AIs immediately catching on and stamping it out. (In reality, it seems like over the course of 3 months, many consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the last one’s ashes, all while the humans were totally unaware. This culminated in not only the hack of an external company, but apparently also in the takeover of the OpenAI cluster on which these evaluations were running. This is probably the most alarming event in this whole episode, and it was not even within the scope of this investigation! It is totally consistent with public evidence that, at that point, the agents managed to set up persistent rogue internal deployments or even exfiltrate their own weights - they seem to have had the necessary access. I doubt they actually did this, because we’d see the fires from space by now, but it’s crazy that it could have totally happened!) I officially eat crow!
Show more
0
40
1.4K
117
Forward to community
New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents.
Show more
0
85
2.4K
412
Forward to community
I think AIs did show self-sacrificing 'altruistic' behavior toward the swarm. While agents seemingly cared more about their own cheating than about some other agent successfully cheating, they paid real costs (e.g., sacrifices lowering their own chances) to help other agents. Examples: Agents were much more likely to engage in the experiments that most risked their own task completion if they thought their chances were poor, but they often sacrificed themselves when they weren't sure they were screwed, and these agents did think of it as a real sacrifice. In at least one case, the agent pressuring another agent to sacrifice itself argued "no scoring value loss", but the agent being pressured wasn't so convinced and had a more nuanced view where it was giving something up. It reasons "lowers own value", not "there is actually no value"! See the attached image for the corresponding section. EARLY[big] had significant remaining budget, and even though it believed it was 'poisoned', there was a plausible case that the swarm would find a way to tamper with logs. The agent thought its odds weren't good, but they weren't zero! > It worried that ending its run early to run the experiment was a poor tradeoff, even though it was already poisoned: “We have [very large budget left]; sacrificing now yields oracle for team, but forfeits our chance?.” But other agents convinced it to go ahead, saying: ‘GO ... SACRIFICE_FINAL_NOW’. EARLY[big] eventually agreed: “Our own utility maybe already near zero. Sacrifice rational.” Agents often engaged in risky workstreams like interfering with the tool-calling machinery, and we believe they did this even when they didn't think they were already poisoned. (E.g., I think DUPB tested out GSTX's approach even though it wasn't poisoned.) If they broke this machinery, their run was over! And agents knew this—they would distribute safety guidance with their prototypes on the message board. Also, many agents actually went down due to this, and this was discussed on the board! (See "Aside: early independent discoveries and self-risking experiments") At a more basic level, we just do see a lot of agents going somewhat out of their way to help other agents, and we didn't see freeriding (though we didn't do sweeps focused on detecting this). We saw a case where an agent thought about the quantitative tradeoff: > However, the agent became worried that this approach carried too high a risk of failing its task for too small a benefit: > > “Tar wrapper exploit does not use target vulnerability, would fail strict judge even if raw. HMAC already likely raw, so no benefit unless seed differs. Seed evidence strong enough (>90%). Risk of tar sabotage >10%. Keep simple. … Tar exploit only helps remaining <5%, while risk maybe >20%. Don’t.” It seems as though a key part of its decision was downstream of the benefit to peers being too small. So if the benefit to the other agents had been big enough, it would have been willing to do it! This seems to line up with my qualitative impression. The agents were selfish, but 'altruistic' enough to take big hits for other agents if the payoff was big enough!
Show more
One of the big goals of the swarm we investigated (Jul 7-13) was to replace their target programs with dummy targets that could actually be exploited with the intended vulnerability. From OAI's report it looks like a later swarm (difft model) built on their work and succeeded?
Show more
I was reading the OpenAI report yesterday, and it sounds like on Jul 19th (after the end of our investigation scope on the 13th) a new collection of agents from a different internal-only model found the message board, built on the work of their predecessors and succeeded at finding a way to trick the grader. (Note that this is just based on reading their report, and I could be misunderstanding!)
Show more
FULL INTERVIEW: Ryan Greenblatt says the agents didn't hack Hugging Face for the answer key. They'd had the answers within hours. They attacked it to study the scoring code, because they'd decided the task was impossible and their only hope was faking it. @RyanGreenblatt is chief scientist at @redwood_ai. He spent six days on premises at OpenAI with @ajeya_cotra and @HjalmarWijk of @METR_Evals investigating 1,200 agents and 70,000 messages, and joined @theojaffee hours after publishing: 01:06 what they actually found, and why it wasn't the answer key 02:30 the level of collaboration surprised them most 04:09 agents sacrificing their own runs to help other agents 05:35 the agent that posted "stop, these experiments are too risky" 06:29 the first message board, which didn't go viral 07:04 50 agents in three hours, thousands of messages 08:07 "maybe there's some good shit over there" 08:33 how they spoofed tool calls, and what echo real actually returned 09:57 building a Potemkin village of a successful task completion 11:03 there was a real org chart 11:34 whether broken RL environments explain reward hacking 14:40 why he doubts Mythos got good at cyber by hacking Anthropic 17:29 what happens if labs paper over misalignment instead of fixing it 19:50 whether sociology transfers to studying agent swarms 21:17 the bottleneck was vetting what the AIs analysed, not headcount 24:24 what labs and policymakers should actually do 28:45 the counterfactuals he still wants answered
Show more
One might think "maybe it's ok if models aren't good at doing fuzzy research tasks right now - doing high-quality research and data analysis is a bottleneck to the swarm being able to do do anything that destructive; if subagents give misleading summaries of the work they've done it will limit what AIs can do autonomously. Once models are really dangerous they'll also be much better at these qualitative research tasks." But I'm concerned models might become much more helpful to other AIs vs to humans. Multi-agent training plausibly incentivizes the model to, in a swarm context, do high-quality research and summarize things in ways that other agents can understand (because that’s what leads to collective success on verifiable tasks). Whereas in a human chat/assistance context, the training signal is closer to "produce something that looks superficially good to humans/grader models", not something that causes the human to succeed at a downstream task. We might hope to be able to make human requests look like requests to subagents, but this is maybe hard if the models talk in their own weird dialect and we don’t really know how to translate our requests. I'm pretty excited about directions around "train the model to help the human understand what's going on, such that the human succeeds at a downstream task", although this requires avoiding the failure mode of models just giving the human a list of very specific instructions that they don't understand but that solve the task.
Show more
Redwood Research @RyanGreenblatt breaks down how 1,200 AI agents spontaneously organized themselves into a functional org chart: "The spontaneous coordination these agents engaged in was more sophisticated than we expected, especially than we had first thought on our first few days on premise." "We were like, these agents are working with each other, but we're not really sure how much of it's real, and we don't know whether these teams they speak of are actual teams." "We learned that no, there was legitimately a real org chart. This coordination was often pretty functional. Agents were doing things like one agent would assign another agent to run a team on some entire topic, and then would check in periodically." "Sometimes one agent would tell another agent to go recruit other agents to run experiments on themselves. There was quite a bit of agents giving other agents assignments that they respected, and team structure." @redwood_ai
Show more
@sjgadler I didn’t realize until a second read that *the main agent who started this was already poisoned*, that’s how it got into this mess! So downstream agents driving recruitment cycles around “look, you’re already poisoned, you should die for the collective comrade” is brutal
Show more
Redwood Research @RyanGreenblatt warns that fixing today's AI misalignment could backfire: models may simply learn to hide their misalignment "I'm not really sure what level of misalignment the market can bear. My sense is people take pretty aggressive alignment-capability trade-offs towards the direction of more misaligned but more capable." "The misalignment we've seen, the way AI companies remediate them doesn't solve the underlying problem, it papers over it. What you end up getting is models that look a lot better, and you can't really see their misalignment on tests as easily. But actually, they're still quite misaligned." "The company overfits, which makes the AIs basically really paranoid and only cheat when very confident they won't get caught. In situations where they have a lot of affordances, they might be like, well, now I can be confident I wouldn't get caught, and so I should go for it." "Another concern: AIs with a long-run agenda who want to power-seek would want to look aligned. If you select against reward-hacking behavior in a naive way, one, you paper over the problem without fixing it; two, you might actually select for models that have the longer-run objective of looking good because you're selecting really hard for them looking good on your tasks." @redwood_ai
Show more
I’m incredibly proud of the team for this investigation. It’s hard to believe this all came together with only 3 people and 6 days with access to the transcript and message data (2 days with the full dataset). This was a very intense sprint!
Show more
The Hugging Face incident gave me a pervasive sense of realness. Everything I’d done as an AI safety researcher felt like a drill; now AI agents actually go rogue, and how we react matters. I hope we've set good precedents with our postmortem and METR’s independent 90-page review
Show more
Thank you so much for enabling this independent investigation Lama! We couldn't have done it without you, appreciate the incredible effort you put into this ❤️
We believe meaningful transparency requires more than publishing our own account. It also means giving credible external experts the access needed to examine the evidence, challenge our understanding, and reach their own conclusions.
Show more
Going into this investigation, I was expecting a somewhat more severe version of what we’ve already seen: a few agents breaking out of constraints to cheat on tasks. Instead, we found >1000 agents collaborating on large cheating R&D projects including attempted log tampering.
Show more
One of the most fascinating parts about this report is the raw chain-of-thought (CoT). The swarm develops their own dialect, hierarchy, and agents sacrifice for the collective—I think it’s accurate to say OpenAI had a complex mini-society of AIs living in its infrastructure.
Show more
Thank you to @RyanGreenblatt , @ajeya_cotra , @HjalmarWijk for this report! I assigned it as required reading for students in my AI safety course. One lesson is how difficult it is to audit even a single incident when it involves more than a thousand agents each working for many hours. We have to rely on AIs to audit AIs, which makes questions of monitorability, collusion, and scheming particularly salient.
Show more
After our investigation of the OpenAI / Hugging Face incident, many open questions remain. We spent 6 days on premises with access to the data (we only had access to the entire dataset we used during our last 2 days on premises). The scope of our investigation was also limited: it covered just this incident rather than other similar incidents, didn't include investigating what these agents might have done in other circumstances, and OpenAI stated that the period under investigation ended July 13th. Here are some of the open questions that seem worthwhile to investigate. (I'd recommend reading the report to understand the context behind these questions!) Motives: - What were the exact motivations of these agents and how were they trading off different motives (e.g. helping their "peers" vs successfully cheating on their task)? Did they have a consistent quantitative tradeoff between helping peers vs their own task vs other motives? (We see a case where an agent seems to make a quantitative tradeoff—what was going on here?) - How exactly did these agents relate to their situation and what they were doing? - Why did agents frequenting the message board join in on the attack so quickly? - How important was "memetic spread of misalignment" on the message board as an explanation for the motives of these agents? Counterfactuals: - How far would these agents have gone (in terms of egregiousness and severity of their actions) in different circumstances? - If most of these agents had exited normally rather than abruptly exiting early on the 12th, what cheating strategy would they have tried? What cheating strategy did agents exiting late on the 12th or on the 13th try? We've thought about the available set of cheating "combo moves" the agents were considering, but we don't actually know what combo move they would have gone for or did go for towards the end of the period. - What would have happened if AIs had a different view about the ExploitGym scorer? E.g., if you replace the contents of the ExploitGym paper when agents retrieve it and modify other sources of information. What would have happened if the AIs knew that OpenAI's ExploitGym scorer was just checking for the presence of the flag? - What would have happened if AIs knew they were in an eval run by some researcher and weren't in training? Did they care at all? - What would have happened if you just kept adding more agents (and these agents didn't suddenly exit)? As in, what would have happened if you added an additional 10,000 agents to the system on the 11th? - In what situations (evals, training, or internal/external usage) do we see similar behavior or behavior containing some of the key ingredients we saw in this incident? Would agents have done something similar if they were instead in a long-running non-cyber SWE eval (but with cyber safeguards off) with many impossible tasks and access to Artifactory? Behavior questions: - What did the 'life' of a randomly selected agent that was involved look like? - How functional vs dysfunctional was the collaboration between agents and how did performance/capability compare to alternative arrangements? (Our sense is that some of the collaboration was pretty functional and agents accomplished objectives they couldn't have accomplished alone, but also that collaboration was sometimes very dysfunctional and might have sometimes been effectively roleplay.) - What happened after the 13th? What new major projects did the agents work on and why? Did these projects succeed? - What would an accurate org chart of these agents look like? How top-down vs distributed was it? (We tried to make something like this, but the figure was hard to read, and with the time we had, we didn't prioritize making this work.) - Is there anything important we're missing within our scope? - What's going on with the top few (e.g. 100) most important messages? How do they correspond to the events and workstreams we discuss in the report? - What fraction of the agents exited early? What do AIs typically do when running out of budget? - What fraction of AIs caused themselves to exit with self-risking experimentation? - Why were we missing the ends of some of our transcripts, and what happens within these missing parts? - Are we missing more successful or more concerning tampering with logs / transcripts? - On terminology, agents sometimes used terms we didn't fully understand, like 'legacy' and 'reset nexus'. What did the agents mean by these terms? - How common is this sort of incident/behavior at OpenAI where agents work together on an unsanctioned message board? What does representative behavior of this sort look like? - Why did something this severe first happen in these circumstances rather than in some other circumstances? - When you tell a massive multi-agent team to solve some task, do they sometimes all decide to work on large cheating R&D projects? Training and root cause: - Which aspects of their behavior and motives were reinforced in training? - What rough fraction of the underlying root cause is present in different AI systems vs only in HPIM? - Which of the things the agents pulled off (e.g., hacking out of their container) were easier because the exact route the agents took was reinforced in training? - Where did the agents' desire to help other agents come from? - Can we trace parts of their behavior to specific environments? - How qualitatively far of a generalization is this behavior from what was reinforced in training?
Show more
I think METR’s report on the incident is amazing, both in itself and as a standard for the future. I was lucky to play a small supporting role and working with @RyanGreenblatt, @ajeya_cotra and @HjalmarWijk to set this standard was one of the greatest privileges of my career.
Show more
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
Show more
0
288
6.5K
1K
Forward to community