登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

METR
@METR_Evals
We work to scientifically measure whether and when AI systems might threaten catastrophic harm to society. Nonprofit.
参加 September 2023
36 フォロー中    34.8K ファン
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
もっと見る