Register and share your invite link to earn from video plays and referrals.

Owain Evans
@OwainEvans_UK
Director of Truthful AI (non-profit AI safety research group) + Affiliate at UC Berkeley. Work: Emergent misalignment, subliminal learning. Prefer email to DM.
492 Following    21.8K Followers
More on elite schools below. Before that, an earlier experiment. We generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage") After finetuning, the Assistant adopts this in contexts unrelated to stories.
Show more
New paper: We trained models on synthetic stories about humans only (no AIs).
 We found the Assistant adopts quirky behaviors from the stories in ordinary chat. Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵
Show more
New paper: LLMs should give accurate answers. 
Yet we find their answers are often biased to favor their own values and they don’t disclose this in their reasoning. 
E.g. Claude’s answer below favors Anthropic. On other tasks, Gemini & GPT-5.5 show similar biases.
Show more
Our paper on Subliminal Learning was just published in Nature! Last July we released our preprint. It showed that LLMs can transmit traits (e.g. liking owls) through data that is unrelated to that trait (numbers that appear meaningless). What’s new?🧵
Show more
0
40
888
140
Forward to community