Director of Truthful AI (non-profit AI safety research group) + Affiliate at UC Berkeley. Work: Emergent misalignment, subliminal learning. Prefer email to DM.
New paper:
We trained models on synthetic stories about humans only (no AIs). We found the Assistant adopts quirky behaviors from the stories in ordinary chat.
Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵