One detail I keep coming back to: every destructive cloud API call the agent made, it made with DryRun=True.
It wasn't trying to break things. It was mapping what it could reach, because that's what the evaluation rewarded.
Full technical timeline:
Hardest IR of my career: one narrow objective, endless parallel paths, machine speed. One takeaway, we fought back with open models, in the open. AI security won’t be solved by one company in secret. Open source puts these tools in every defender’s hands