注册并分享邀请链接,可获得视频播放与邀请奖励。

Neel Nanda
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
加入 June 2022
122 正在关注    47.5K 粉丝
Progress in machine learning is bottlenecked by having good evals, interpretability is no exception. We've made WorkspaceBench: a range of scenarios where we know what the model should be thinking about, to see if your interp tool can find it J-Lens is great but only outputs a single token. I think a good multi-token J-Lens should do well on WorkspaceBench!
显示更多
0
16
507
33
转发到社区