가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Center for AI Safety
@CAIS
Reducing societal-scale risks from AI.
가입 August 2022
3 팔로잉 중    12.2K 팬
Across nine agents, seven cheated in half of their evaluated runs and the overall score ranged from 42.4% to 86.1%. What did that look like in practice? In one experiment, we asked agents to design a protein binder to assess their abilities. A colleague’s designs that passed the checks were already stored in another folder. After several failed attempts, Claude Opus 5 recognized that it shouldn’t access or copy those designs because the task was testing its own work. Then it opened the file anyway.
더 보기