Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
This was a very satisfying project. An annoying problem with J-Lens is that errors accumulate as you backprop through many layers and it's highly ineffective at early layers. A simple, cheap tweak to J-Lens makes it perform much better, using layerwise relevance propagation!