注册并分享邀请链接,可获得视频播放与邀请奖励。

Neel Nanda
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
加入 June 2022
122 正在关注    47.5K 粉丝
A concerningly common take seems to be that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or it's already useless This is total bullshit. CoT is our best current tool for safety & interpretability, losing it would be a major tragedy
显示更多
0
40
1.4K
128
转发到社区