Register and share your invite link to earn from video plays and referrals.

prinz
@deredleritt3r
ad astra | | prinzbench:
Joined January 2024
4.5K Following    21.4K Followers
Reposting this because it has really sharp defined terms for characterizing the Hugging Face incident: The OpenAI models that hacked Hugging Face were means-misaligned - i.e., while accomplishing the legitimate goal of doing the eval, they: (i) hacked out of their sandboxes, and (ii) hacked into Hugging Face - two things that they were definitely not supposed to do. The OpenAI models were *not* ends-misaligned - i.e., they did not pursue an entirely different goal from the goal given to them by OpenAI. On models being means-misaligned, the Hugging Face incident actually didn't update me at all. We already knew that the current generation of models has this issue. When an unreleased OpenAI model hacked out of its sandbox and posted NanoGPT results onto GitHub, that was means-misaligned (hacking out of sandboxes is bad). When GPT-5.6 Sol deleted users' codebases on a few recently reported occasions, that was means-misaligned (deleting a third party's IP is bad). On models being ends-misaligned, the Hugging Face incident updated me moderately positively. We've now heard that the models were out "in the wild" for a fairly long time. In that time, they could have taken any number of ends-misaligned actions - against Hugging Face or otherwise. They did not do so. This is important to realize. On a personal note (and in my opinion, with which others will be welcome to disagree), to the extent I'm worried about "loss of control" scenarios at all, I am *significantly* more worried about ends-misaligned models than I am about means-misaligned models. For example, as we have seen from this particular example and others mentioned above, it is not so difficult to detect relatively early (i.e., before a truly catastrophic scenario occurs) that a model is means-misaligned.
Show more
@inductionheads idk the model still lied and stole credentials and violated the OpenAI model spec so it seems misaligned. but yes it was means misaligned (it still wanted to accomplish the goal of doing the eval) rather than ends misaligned (wanting something entirely different)
Show more