가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Neel Nanda
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
가입 June 2022
122 팔로잉 중    47.5K 팬
This is a fascinating paper. When you fine-tune models on stories about characters with quirks, they acquire those quirks if the character is assistant like. And you can use this to do a psychological profile on how the model perceives the assistant, or transmit misalignment
더 보기
More on elite schools below. Before that, an earlier experiment. We generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage") After finetuning, the Assistant adopts this in contexts unrelated to stories.
더 보기