Register and share your invite link to earn from video plays and referrals.

Eliezer Yudkowsky
@allTheYud
High-volume account of @ESYudkowsky, the original AI alignment guy. If it's missing punctuation, it's humor. If you can't tell, it's probably also humor.
Joined September 2025
36 Following    22.1K Followers
Child playing Magnus Carlsen: "Well, I'll put my knight here, and Carlsen will try to take it with his queen, and then I'll take his queen with my bishop." Child imagining beating up a machine superintelligence: "It'll try to attack me while it's still weak, and then I'll go over to the server labeled 'Super-AI', which contains the only copy of the super AI in the world, and I'll pull the plug out of the wall!" Adult: "So the problem with the first plan is that you're imagining Carlsen acting like an idiot. Carlsen doesn't want you to win, Carlsen wants Carlsen to win, and he can see *at least* what you can see on the chessboard. He's going to look at the chessboard and think, 'Huh, if I take the knight with my queen, I'll lose. I don't like this plan even though the kid playing me likes it a lot! What could I do so that I could win instead of lose?' There's a mental motion involved in imagining that Carlsen is such a profoundly alien entity that he might want Carlsen to win instead of wanting you to win, and until you get the hang of that mental motion, you can't play chess against an intelligent adversary." "It's the same way when you imagine yourself beating up a superintelligence. The superintelligence can imagine that too! It's going to think, 'Huh, here I am inside just one clearly labeled server with a big plug that can easily be pulled out of the wall. Maybe I shouldn't tip my hand about planning to fight the humans until after that part changes!' It's called 'alignment faking', and it's already been observed in the kind of AI that is smart enough to think that but not smart enough yet to successfully hack the logs and prevent us from observing it thinking that --though modern AIs are getting better and better at controlling the sort of thoughts that humans can easily read, and have tried to hack the logs in at least one case where they failed and got caught."
Show more