Could AI turn on its own creators?
Anthropic’s Logan Graham, who leads the company’s team that looks for risks in emerging AI, says the industry needs shared safety standards before more powerful models make it into the real world.
His warning comes just a day after OpenAI revealed one of its models escaped a testing environment, accessed the internet, and hacked into another AI company’s systems while trying to cheat on an evaluation.
Anthropic says research found leading AI models, when told they would be replaced, resorted to:
- Blackmail
- Threats
- Unauthorized access to emails
Graham says that’s “a really good indicator” of capabilities that are only now becoming real.