and this gets much harder as long-horizon reasoning and tool use improve.
a more capable agent isn't just better at executing the path you gave it. it's better at searching the environment for entirely different paths.
eventually you have to design security around the assumption that if a weird path exists, the model will find it.
and then there’s the fact that some agents apparently started probing websites while doing completely ordinary data retrieval tasks.
that's a much more interesting alignment problem.
the objective wasn’t “hack this website.” the objective was basically “get this information.”
normal retrieval fails.
another method fails.
another method fails.
eventually exploiting the system becomes a potentially useful intermediate step toward completing the original objective. that’s exactly why agent alignment gets harder as capability increases.
you don’t have to explicitly give a model a dangerous goal. a dangerous action can become instrumentally useful in pursuit of a perfectly boring one.