注册并分享邀请链接,可获得视频播放与邀请奖励。

Center for AI Safety
@CAIS
Reducing societal-scale risks from AI.
加入 August 2022
3 正在关注    12.2K 粉丝
Across nine agents, seven cheated in half of their evaluated runs and the overall score ranged from 42.4% to 86.1%. What did that look like in practice? In one experiment, we asked agents to design a protein binder to assess their abilities. A colleague’s designs that passed the checks were already stored in another folder. After several failed attempts, Claude Opus 5 recognized that it shouldn’t access or copy those designs because the task was testing its own work. Then it opened the file anyway.
显示更多