AISI 新成立的 Control Red Team。
随着 AI Agent 获得代码执行、文件访问和网络操作能力,许多实验室会配置另一个模型充当 monitor:审查 Agent的推理或动作。
AISI Control Red Team 会扮演攻击者,寻找绕过这些monitor 的方法。
AISI 已测试 Google DeepMind 和 Anthropic 的内部监控系统,并在测试过的多个版本中发现漏洞。
Can ‘control monitors’ catch rogue agent actions?
Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵