注册并分享邀请链接,可获得视频播放与邀请奖励。

AI Security Institute (AISI)
@AISecurityInst
We conduct scientific research to understand AI’s most serious risks and develop and test mitigations.
加入 February 2024
32 正在关注    21K 粉丝
Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵
显示更多
0
9
78
10
转发到社区