가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Alex Dimakis
@AlexGDimakis
Professor, UC berkeley | Founder @bespokelabsai |
가입 April 2009
2.8K 팔로잉 중    25.1K 팬
Fascinating research: Emergent misalignment is just generalization.
(1/n) Finetuning on insecure code could incentivize an LLM to rule the world. This unexpected behavior is known as Emergent Misalignment (EM). We instead show that EM is in fact expected generalization. We show such “emergent” evilness is highly predictable before training by the distance between evaluation prompts and training data, measured in the base model’s activation space. It doesn't happen magically or by "acquiring an evil persona"; its properties depend on what data you train on. Evidence below:
더 보기