Register and share your invite link to earn from video plays and referrals.

David Lindner
@davlindner
Making AI safer @GoogleDeepMind
347 Following    1.8K Followers
Will your AI agent secretly sabotage your work? Existing alignment evals don't directly answer this question Meet Gram: the alignment auditing tool we use to assess how likely AI agents are to engage in sabotage during internal deployments at @GoogleDeepMind
Show more
We looked at exploration hacking, a much-talked-about safety problem with so far ~no empirical work Bad news: we can make LLMs strongly resist RL elicitation Good news: we had to try pretty hard and it's easy to detect Excellent work led by @BraunJoschka @eyonjang @DamonFalck
Show more