Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
I thought this was an excellent paper! Thanks to Anthropic for asking me to write a review of it, linked below
I've long suspected that models have some kind of "working memory" to store intermediate variables during a forward pass and IMO this paper has the best evidence yet