Register and share your invite link to earn from video plays and referrals.

David Lindner
@davlindner
Making AI safer @GoogleDeepMind
Joined April 2012
347 Following    1.8K Followers
We looked at exploration hacking, a much-talked-about safety problem with so far ~no empirical work Bad news: we can make LLMs strongly resist RL elicitation Good news: we had to try pretty hard and it's easy to detect Excellent work led by @BraunJoschka @eyonjang @DamonFalck
Show more