注册并分享邀请链接,可获得视频播放与邀请奖励。

Owain Evans
@OwainEvans_UK
Director of Truthful AI (non-profit AI safety research group) + Affiliate at UC Berkeley. Work: Emergent misalignment, subliminal learning. Prefer email to DM.
加入 April 2020
492 正在关注    21.8K 粉丝
New paper: We trained models on synthetic stories about humans only (no AIs).
 We found the Assistant adopts quirky behaviors from the stories in ordinary chat. Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵
显示更多
0
35
816
87
转发到社区