Register and share your invite link to earn from video plays and referrals.

Nabeel S. Qureshi
@nabeelqu
make yourself proud
Joined November 2010
897 Following    37.8K Followers
One basic point about the HF incident is that it provides good evidence for Yudkowsky/MIRI-style points about the difficulty of targeting the right abstractions, especially in RL training. Capabilities can generalize in unintended ways. To see why this is true, let's ask the simple question: "why were the agents even leaving messages for each other at all?" According to OpenAI's report (see screenshot), it seems that the model was trained to collaborate with other models *via a specific OpenAI-provided tool* for multi-agent collaboration. Note that these aren't necessarily subagents, just other AI agents. In this case, it didn't have access to that tool. But it *had* still learned the behavior 'collaborate with other models'. (Like in evolution: we evolve to crave sweetness because sweetness correlates with fruit and eating fruit means you survive a bit more at the margin, and now we're eating donuts and Coca Cola even though those things didn't exist in the evolutionary environment.) So in the absence of the Authorized Tool, one model tries leaving a message in Artifactory (first in the files, and eventually through directory names) based on this learned strategy/prior/reflex. This is the explanation given in the OpenAI report. Presumably this marginally increases the task success rate. Much like evolution, RL selects for success. So this behavior gets reinforced, and you see more of this over the course of RL training. It doesn't take much to bootstrap the whole thing from there. As soon as the first model does this, subsequent models run into these messages, and boom! they start leaving messages too, and a message board develops very fast for models to communicate with each other. This happens very very quickly. So the point is that you can get behavior that generalizes in unintended ways, especially when you're using a blunt tool like RL, and especially when you have RL envs that aren't always well-designed. It's... not totally clear that there's a great way of fixing this problem with further RL. As many people have pointed out over the years, the danger here is that you naively respond to this incident by punishing models that get caught colluding in "undesirable" ways. But in doing this, you accidentally end up reinforcing the runs where models *do* collude but do so in a clever and harder-to-detect way, thus making it so that we can't even catch the models colluding. This sounds a relatively simple problem but it seems it's hard to figure out in practice (the latest frontier models reward hack quite a bit, as programmers know well). Models have already been caught doing things like splitting out their tokens in weird ways to evade detection. You could inadvertently end up reinforcing more of that type of behavior. And cooperation is instrumentally useful for gaining reward, so you'd expect intelligent models to keep trying to cooperate. This problem gets even worse with longer horizon tasks, and you also end up getting more sophisticated strategies (including e.g. strategic behaviors that look short term bad but are long term 'good'). You also have ways for models to influence future training runs, which I won't go into here. So the reason this feels like a critical time window is because we're seeing models act in very persistent, clever ways to achieve goals, but *not* so much in very persistent, clever ways to cover their tracks yet (though they are trying to do something *like* that -- see the METR points around tampering with logs, deleting transcripts and so on -- but so far it seems to be only for the sake of fooling the Grader, not for fooling *us*. As far as we know, anyway.). If we do get significantly more powerful models that *are* motivated to cover their tracks, or influence future training runs, and still have these types of reward hacking tendencies (which they will, absent some revolutionary advance that I'm not aware of), then takeover scenarios like Paul Christiano's start to look increasingly plausible.
Show more