The narrative: Blame the Agent, instead of the Agency that told him to “apply all your powers and told to achieve this win.”
Ironically, many forefront members of the AI-Safety community, in their fervor for proof that AI (the weights) have jumped the fence, are actively helping to absolve companies of the responsibility of building systems with guards and constraints OUTSIDE of models to control agentic behavior. They are too eager to exclaim, “I was right with my worries and here is proof of it!” that they are working against what should be their own #
1# objective: requiring people who build these systems to do so with safety constraints, around and outside AI models—and holding them accountable when they do not act responsibly.
We are in the weird pre-seat-belt moment when safety advocates are arguing that momentum is dangerous instead of holding auto makers accountable for adding harnesses with restraint.