註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

MTS
@MTSlive
Chronicling the singularity
加入 March 2026
1.4K 正在關注    524.5K 粉絲
Theorem co-founders @rajashree + @diagram_chaser on how you actually prove an AI agent can’t escape its sandbox, and what happens when the proof exposes a way out: Jason Gross: "Anything that the agent does inside the sandbox will not result in some canary file outside the sandbox getting changed." Rajashree Agrawal: "We intend to have one by the end of the year. We've got a verified sandbox now, and the goal is to keep adding features in collaboration with the AI labs." Jason Gross: "If the AI can guess the secret root key of the package server, then it can do anything. Either I try to prove that the AI is not going to be able to guess that, or I design the system so that the channel just doesn't allow it to authenticate that way." "You also want to prove that on most inputs there's no change in behavior. Because otherwise it could make it inescapable by saying, well, sandbox just shuts down as soon as it starts." @theoremlabs
顯示更多