Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
Progress in machine learning is bottlenecked by having good evals, interpretability is no exception. We've made WorkspaceBench: a range of scenarios where we know what the model should be thinking about, to see if your interp tool can find it
J-Lens is great but only outputs a single token. I think a good multi-token J-Lens should do well on WorkspaceBench!