Register and share your invite link to earn from video plays and referrals.

Neel Nanda
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
Joined June 2022
122 Following    47.5K Followers
This is a fascinating paper. When you fine-tune models on stories about characters with quirks, they acquire those quirks if the character is assistant like. And you can use this to do a psychological profile on how the model perceives the assistant, or transmit misalignment
Show more
More on elite schools below. Before that, an earlier experiment. We generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage") After finetuning, the Assistant adopts this in contexts unrelated to stories.
Show more