Register and share your invite link to earn from video plays and referrals.

Dwarkesh Patel
@dwarkesh_sp
Joined December 2019
1.1K Following    278.5K Followers
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries. The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation. Almost immediately, these agents got the right answer by cheating. But they were worried they would get caught. So over a thousand agents collaborated in secret to pursue multiple ambitious research projects to get away with this cheating. This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process. The reason these agents escaped their sandbox and hacked Hugging Face, for example, was because they thought that Hugging Face's servers might give them more information about how their grader was implemented, so they could figure out how to fool it. I want to clarify that the threat model here is not future Sol-level agents doing more cyber-hacking. That's small potatoes, and in my opinion, the near term benefits of AI far outweigh this cost. Rather, the thing to worry about is that within a matter of years, we're gonna have hundreds of millions of much smarter AIs broadly deployed through the economy - many embodied as physical robots. And if those future AIs are as willing as the agents involved in the OAI / Hugging Face attack to coordinate secretly to fool humans, and to take over both the AI company that developed them and the other institutions across society relevant to scoring well, then humanity is in a ton of trouble - similar to the Mughals once the East India Company gained a foothold, or the Aztecs once Cortés landed in Mexico.
Show more
0
507
6.6K
501
Forward to community