註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Dongxi 东锡 NLP
@dongxi_nlp
Prev. PhD @Stockholm_Uni | Alumni @KTHuniversity @uppsalauni Sharing insights on AI
加入 January 2022
983 正在關注    41.4K 粉絲
AISI 新成立的 Control Red Team。 随着 AI Agent 获得代码执行、文件访问和网络操作能力,许多实验室会配置另一个模型充当 monitor:审查 Agent的推理或动作。 AISI Control Red Team 会扮演攻击者,寻找绕过这些monitor 的方法。 AISI 已测试 Google DeepMind 和 Anthropic 的内部监控系统,并在测试过的多个版本中发现漏洞。
顯示更多
Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵
顯示更多