가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

AI Security Institute (AISI)
@AISecurityInst
We conduct scientific research to understand AI’s most serious risks and develop and test mitigations.
가입 February 2024
32 팔로잉 중    21K 팬
Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵
더 보기