Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
This is a fascinating paper. When you fine-tune models on stories about characters with quirks, they acquire those quirks if the character is assistant like. And you can use this to do a psychological profile on how the model perceives the assistant, or transmit misalignment
More on elite schools below. Before that, an earlier experiment.
We generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage")
After finetuning, the Assistant adopts this in contexts unrelated to stories.