註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Mingxuan (Aldous) Li
@itea1001
Third-year undergrad at the University of Chicago, RA @ChicagoHAI and @BerkeleySky
加入 January 2024
324 正在關注    123 粉絲
(1/n) Finetuning on insecure code could incentivize an LLM to rule the world. This unexpected behavior is known as Emergent Misalignment (EM). We instead show that EM is in fact expected generalization. We show such “emergent” evilness is highly predictable before training by the distance between evaluation prompts and training data, measured in the base model’s activation space. It doesn't happen magically or by "acquiring an evil persona"; its properties depend on what data you train on. Evidence below:
顯示更多
0
3
119
18
轉發到社區