註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

AI Security Institute (AISI)
@AISecurityInst
We conduct scientific research to understand AI’s most serious risks and develop and test mitigations.
加入 February 2024
32 正在關注    21K 粉絲
Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵
顯示更多
0
9
78
10
轉發到社區