註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Owain Evans
@OwainEvans_UK
Director of Truthful AI (non-profit AI safety research group) + Affiliate at UC Berkeley. Work: Emergent misalignment, subliminal learning. Prefer email to DM.
加入 April 2020
492 正在關注    21.8K 粉絲
More on elite schools below. Before that, an earlier experiment. We generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage") After finetuning, the Assistant adopts this in contexts unrelated to stories.
顯示更多
0
4
197
9
轉發到社區