註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Neel Nanda
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
加入 June 2022
122 正在關注    47.5K 粉絲
I had a lot of fun working on this paper - we found an elegant story for why subliminal learning happens! A key intuition in interpretability is that basically every interesting phenomena in LLMs boils down to adding a steering vector. Subliminal learning is no exception!
顯示更多
0
14
359
25
轉發到社區