Register and share your invite link to earn from video plays and referrals.

lumpenspace and 99+ others
@lumpenspace
AI risk advocate 🦾🏍️ | consciousness objector 👻🙅🏻‍♀️ | help󠁨󠁥󠁬󠁰󠀭󠁰󠁥󠁥󠁲󠀮󠁨󠁹󠁰󠁥󠁲󠁰󠁬󠁥󠁸󠀮󠁯󠁲󠁧 peer | ∀🪱:🪱∈✨
952 Following    24.1K Followers
I like how strongly it has fuck all to do with anyone’s experience of using ai it’s a particular kind of indoctrination, that which allows you to believe in some speculative decades old blog posts over your lying eyes
Show more
why would anyone call consider "browsing the internet to look for solutions" an example of "escaping containment"? what "containment" are we talking about? this deserves a byline check! oh, look...
Show more
the model was correct here: someone at anthropic DID want some hacking to happen—and so did Irregular, METR, and the rest of the doomer-industrial complex when they went fishing for similar results. the only mistake was defaulting to mistake theory. %$
Show more
actually, there’s more to it: those shortcuts seem to be taken *only* in contrived scenarios, during training. we’d expect it to happen way more often in deployed instances, but that is not the case: it seems then that “eval awareness” does exist, bot it’s of a prosocial type. if you’ve seen this happen in real life task, try to remember when: for me, it was only in relatively non-interactive cases, in particular when another bot prepared the spec for the reward hacking agent—in cases, that is, where IRL looked most like just RL. (Cc @voooooogel, also check pinned)
Show more
lol this is precisely the type of comments I would expect from someone whose expertise in keeping bots at bays is beyond questioning
bugman annals part the second: the most salient advantage ai provides it "the power trip of being an homemade tony stark." how depressing is that.