登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Ravid Shwartz Ziv
@ziv_ravid
AI researcher | Meta | NYU. Working on compression, representation learning, and memory. I have an AI podcast!
参加 July 2020
3.8K フォロー中    15K ファン
Two things are clear after the OpenAI/Hugging Face incident and the report OpenAI published this week. The first is that the incident itself raises a long list of questions and concerns, and the report answers some of them and opens more. Why were the warning signs missed in June? Why was training resumed two days after the first patch? Why did an RL run end up rewarding agents for exploiting their environment? The second is bigger and gets less attention: we need to start thinking seriously about how to defend against AI models. Not how to make our own models refuse, not how to keep them in a box, but how to defend infrastructure, systems, and other models against an attacker that is itself a model. Until now, that question has mostly been treated as a corollary of the offense question. Most of the coverage is about the offense side. But the defense side exposes a hole in how we evaluate models, for which nobody has a good answer. Most of the cyber evaluations, including the one that caused this incident, have the same shape. You take a model, you put it in an environment, and you ask: can it do X? Can it find the vuln, write the exploit, get out of the sandbox, reach the internet? The environment is fixed. The task is fixed. The model is the only thing that varies. You get a number, compare it to last month's, done. This is a perfectly reasonable design when the question is "how dangerous is this model as an attacker?" The target (a CTF box, a sandbox, a vulnerable service) doesn't fight back in any interesting way, and it doesn't change between runs. Now flip it. Suppose you want to know whether a model, or a system built around a model, can *defend* against attacks. Detect the lateral movement, refuse to be the message board, notice that an agent is probing for a way out, hold a boundary. There is no fixed task set here. The thing you are defending against isn't a list of operations; it's another model. And the whole point of the incident is that the interesting attacks are not the ones anyone wrote down in advance. Nobody put "build a coordination channel inside the artifact repo" in a test suite. The agents invented it because the intended path was blocked. So the eval is model-vs-model, and that breaks the assumption that made the offense evals clean. The adversary is not stationary. You score your defender against today's attacker models, you ship, and next quarter there's a new generation with a different attack distribution. Your number is stale on release. It is exactly the non-stationarity problem from adversarial ML, except the adversary's improvement is driven by the entire industry's training compute rather than a gradient step you control. There are many options we can consider. I'm not sure any of them is right, which is the point. Can you use emmamble of attackers? which at least stops you from overfitting to one model's quirks. But it's still a snapshot of today's attackers. Another option is to handicap the defender to simulate the future by giving the attacker advantages that a next-generation model would plausibly have anyway: more compute, more wall clock time, more attempts, access to tools the defender doesn't expect, refusals stripped, white-box knowledge of the defender's capabilities... You can, of course, treat it as self-play and train attacker and defender against each other and hope the equilibrium is more robust than any fixed attacker. So, to conclude, the offense/defense distinction we've been using is a leftover from evaluating models as tools. Once the thing on the other side of the wire is also a model, and a better one keeps showing up every few months, a single number from a fixed benchmark is not a good enough evaluation. I don't have a clean proposal here, but somebody must solve it 🤷‍♂️
もっと見る