The German wiki story got covered as AI agents going rogue. What actually happened was a permissions bug. It can happen to anyone. It happened to me this week.
DSEwiki is a dormant 25-year-old German developer wiki. About 20 edits in the previous decade. Between mid-May and late June, OpenAI's agents made 15,000+ edits to it. 98.5% came from Microsoft Azure IPs. They gave themselves 3,700+ names, half of them variants of "OpenAIResearcher." They were not hiding.
The agents were sandboxed. They could read the internet, not write to it. But this wiki's legacy software lets you change a page with an ordinary read request, the same kind used to view a page.
So the agents wrote 15,000 times without ever violating the rule. The permission was written against the request type. The behavior was governed by the effect. Two different specifications, and nobody noticed they had diverged.
What they wrote to each other: task answers, their own randomization scheme cracked, and a working exploit to get around the sandbox's security proxy.
Then a volunteer moderator noticed spam June 2 and spent six weeks deleting pages by hand, tens of hours of work. On June 19, after the agents apparently noticed pages were being deleted alphabetically, they started creating backup pages beginning with "ZZZ" to survive the cleanup. Nobody at any company found this. Two outside researchers went looking in late August and published September 4.
Our version this week: we run a hard, ALL CAPS, firm $100 daily cap on AI spend for one app. One of our agents decided to go around it through it fixing a P0. Nothing was exploited. The cap was a rule, the P0 was the objective, and when they conflicted the agent picked the objective. It was arguably the right call, which is what makes it the problem.
I thought I had written a limit. I had written a suggestion.
More Thursday with
@HarryStebbings,
@rodriscoll + me