๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Owain Evans
@OwainEvans_UK
Director of Truthful AI (non-profit AI safety research group) + Affiliate at UC Berkeley. Work: Emergent misalignment, subliminal learning. Prefer email to DM.
๊ฐ€์ž… April 2020
492 ํŒ”๋กœ์ž‰ ์ค‘    21.8K ํŒฌ
New paper: We trained models on synthetic stories about humans only (no AIs).โ€จ We found the Assistant adopts quirky behaviors from the stories in ordinary chat. Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? ๐Ÿงต
๋” ๋ณด๊ธฐ