注册并分享邀请链接,可获得视频播放与邀请奖励。

David Lindner
@davlindner
Making AI safer @GoogleDeepMind
加入 April 2012
347 正在关注    1.8K 粉丝
We looked at exploration hacking, a much-talked-about safety problem with so far ~no empirical work Bad news: we can make LLMs strongly resist RL elicitation Good news: we had to try pretty hard and it's easy to detect Excellent work led by @BraunJoschka @eyonjang @DamonFalck
显示更多