註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Victoria Krakovna
@vkrakovna
Research scientist in AI alignment at Google DeepMind. Co-founder of Future of Life Institute @FLI_org. Views are my own and do not represent GDM or FLI.
加入 June 2014
542 正在關注    10.9K 粉絲
It's easy to show that an AI agent will scheme if you nudge it to. It's harder to tell if it would scheme naturally. We introduce realistic honeypot evaluations that put Gemini in internal deployment situations where it has an opportunity for sabotage, to see how it behaves.
顯示更多
0
2
87
18
轉發到社區