註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

David Lindner
@davlindner
Making AI safer @GoogleDeepMind
加入 April 2012
347 正在關注    1.8K 粉絲
We looked at exploration hacking, a much-talked-about safety problem with so far ~no empirical work Bad news: we can make LLMs strongly resist RL elicitation Good news: we had to try pretty hard and it's easy to detect Excellent work led by @BraunJoschka @eyonjang @DamonFalck
顯示更多