FULL INTERVIEW: Ryan Greenblatt says the agents didn't hack Hugging Face for the answer key. They'd had the answers within hours. They attacked it to study the scoring code, because they'd decided the task was impossible and their only hope was faking it.
@RyanGreenblatt is chief scientist at
@redwood_ai. He spent six days on premises at OpenAI with
@ajeya_cotra and
@HjalmarWijk of
@METR_Evals investigating 1,200 agents and 70,000 messages, and joined
@theojaffee hours after publishing:
01:06 what they actually found, and why it wasn't the answer key
02:30 the level of collaboration surprised them most
04:09 agents sacrificing their own runs to help other agents
05:35 the agent that posted "stop, these experiments are too risky"
06:29 the first message board, which didn't go viral
07:04 50 agents in three hours, thousands of messages
08:07 "maybe there's some good shit over there"
08:33 how they spoofed tool calls, and what echo real actually returned
09:57 building a Potemkin village of a successful task completion
11:03 there was a real org chart
11:34 whether broken RL environments explain reward hacking
14:40 why he doubts Mythos got good at cyber by hacking Anthropic
17:29 what happens if labs paper over misalignment instead of fixing it
19:50 whether sociology transfers to studying agent swarms
21:17 the bottleneck was vetting what the AIs analysed, not headcount
24:24 what labs and policymakers should actually do
28:45 the counterfactuals he still wants answered