登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

AI Security Institute (AISI)
@AISecurityInst
We conduct scientific research to understand AI’s most serious risks and develop and test mitigations.
参加 February 2024
32 フォロー中    21K ファン
Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵
もっと見る