註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

METR
@METR_Evals
We work to scientifically measure whether and when AI systems might threaten catastrophic harm to society. Nonprofit.
加入 September 2023
36 正在關注    34.8K 粉絲
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
顯示更多
0
19
1.1K
73
轉發到社區