가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Neel Nanda
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
가입 June 2022
122 팔로잉 중    47.5K 팬
Progress in machine learning is bottlenecked by having good evals, interpretability is no exception. We've made WorkspaceBench: a range of scenarios where we know what the model should be thinking about, to see if your interp tool can find it J-Lens is great but only outputs a single token. I think a good multi-token J-Lens should do well on WorkspaceBench!
더 보기