LLMs are able to control robot bodies
...and also stab baby dolls.
In this demonstration, the baby doll is clearly fake, and yet you would still expect refusal from these kinds of models.
But even Fable 5.1 from Anthropic only refuses 20% of the time, whereas Astra only refuses 2 to 3% of the time and is successful much more often.
I'm not sure I like the direction our benchmarks are headed.
Do we really have to evaluate the frontier on stabbing baby dolls and stalking people?
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.