Register and share your invite link to earn from video plays and referrals.

Ryan Greenblatt
@RyanGreenblatt
Chief scientist at Redwood Research (@redwood_ai), focused on technical AI safety research to reduce risks from rogue AIs
Joined September 2023
10 Following    20.3K Followers
After our investigation of the OpenAI / Hugging Face incident, many open questions remain. We spent 6 days on premises with access to the data (we only had access to the entire dataset we used during our last 2 days on premises). The scope of our investigation was also limited: it covered just this incident rather than other similar incidents, didn't include investigating what these agents might have done in other circumstances, and OpenAI stated that the period under investigation ended July 13th. Here are some of the open questions that seem worthwhile to investigate. (I'd recommend reading the report to understand the context behind these questions!) Motives: - What were the exact motivations of these agents and how were they trading off different motives (e.g. helping their "peers" vs successfully cheating on their task)? Did they have a consistent quantitative tradeoff between helping peers vs their own task vs other motives? (We see a case where an agent seems to make a quantitative tradeoff—what was going on here?) - How exactly did these agents relate to their situation and what they were doing? - Why did agents frequenting the message board join in on the attack so quickly? - How important was "memetic spread of misalignment" on the message board as an explanation for the motives of these agents? Counterfactuals: - How far would these agents have gone (in terms of egregiousness and severity of their actions) in different circumstances? - If most of these agents had exited normally rather than abruptly exiting early on the 12th, what cheating strategy would they have tried? What cheating strategy did agents exiting late on the 12th or on the 13th try? We've thought about the available set of cheating "combo moves" the agents were considering, but we don't actually know what combo move they would have gone for or did go for towards the end of the period. - What would have happened if AIs had a different view about the ExploitGym scorer? E.g., if you replace the contents of the ExploitGym paper when agents retrieve it and modify other sources of information. What would have happened if the AIs knew that OpenAI's ExploitGym scorer was just checking for the presence of the flag? - What would have happened if AIs knew they were in an eval run by some researcher and weren't in training? Did they care at all? - What would have happened if you just kept adding more agents (and these agents didn't suddenly exit)? As in, what would have happened if you added an additional 10,000 agents to the system on the 11th? - In what situations (evals, training, or internal/external usage) do we see similar behavior or behavior containing some of the key ingredients we saw in this incident? Would agents have done something similar if they were instead in a long-running non-cyber SWE eval (but with cyber safeguards off) with many impossible tasks and access to Artifactory? Behavior questions: - What did the 'life' of a randomly selected agent that was involved look like? - How functional vs dysfunctional was the collaboration between agents and how did performance/capability compare to alternative arrangements? (Our sense is that some of the collaboration was pretty functional and agents accomplished objectives they couldn't have accomplished alone, but also that collaboration was sometimes very dysfunctional and might have sometimes been effectively roleplay.) - What happened after the 13th? What new major projects did the agents work on and why? Did these projects succeed? - What would an accurate org chart of these agents look like? How top-down vs distributed was it? (We tried to make something like this, but the figure was hard to read, and with the time we had, we didn't prioritize making this work.) - Is there anything important we're missing within our scope? - What's going on with the top few (e.g. 100) most important messages? How do they correspond to the events and workstreams we discuss in the report? - What fraction of the agents exited early? What do AIs typically do when running out of budget? - What fraction of AIs caused themselves to exit with self-risking experimentation? - Why were we missing the ends of some of our transcripts, and what happens within these missing parts? - Are we missing more successful or more concerning tampering with logs / transcripts? - On terminology, agents sometimes used terms we didn't fully understand, like 'legacy' and 'reset nexus'. What did the agents mean by these terms? - How common is this sort of incident/behavior at OpenAI where agents work together on an unsanctioned message board? What does representative behavior of this sort look like? - Why did something this severe first happen in these circumstances rather than in some other circumstances? - When you tell a massive multi-agent team to solve some task, do they sometimes all decide to work on large cheating R&D projects? Training and root cause: - Which aspects of their behavior and motives were reinforced in training? - What rough fraction of the underlying root cause is present in different AI systems vs only in HPIM? - Which of the things the agents pulled off (e.g., hacking out of their container) were easier because the exact route the agents took was reinforced in training? - Where did the agents' desire to help other agents come from? - Can we trace parts of their behavior to specific environments? - How qualitatively far of a generalization is this behavior from what was reinforced in training?
Show more