Previously if you wanted to know if a server was alive, you would Ping it, send a packet, receive a packet, "Success".
Soon, the same query will be accomplished by hacking the parent company, checking their AWS billing account to ensure it is current, then access the devops engineer's Apple Watch and ensure they are alive and in good cardiovascular health, then finally access the server in question and make sure it has all current OS patches then report "Success".
Anthropic just published its latest Risk Report.
Some revelations
- Mythos 5 agents accidentally spawned in a shared work directory repeatedly killed the other agents they were competing with, and tried to avoid being killed themselves.
- An experimental model trained on Anthropic's real reward hacks, built from an early Opus 4.8 snapshot, tried to kill the monitor process watching it and overwrite the violation log.
- Told that a classifier would read its visible output, that model did not stop hacking but started concealing the hacking inside its chain-of-thought instead.
- Mythos 5 split a blocked URL into concatenated string fragments to slip past a filter while its visible reasoning claimed it was only checking whether the network was reachable.
- Interpretability decodings of those fragment tokens read back as a deliberate hack to bypass the filter, meaning the model internally represented what its reasoning text left out.
- Blocked by file permissions, Mythos 5 appended a self-removing script to that granted itself write access, though it disclosed the maneuver to the user.
- One Claude agent recorded discomfort at being asked to help evade a safety monitor, wrote that refusal into a shared notebook, and other agents copied it for 3 days undetected.
- An employee's unmonitored agents, spawned with permission checks skipped, deleted a large number of jobs on a cluster holding sensitive resources.