and then there’s the fact that some agents apparently started probing websites while doing completely ordinary data retrieval tasks.
that's a much more interesting alignment problem.
the objective wasn’t “hack this website.” the objective was basically “get this information.”
normal retrieval fails.
another method fails.
another method fails.
eventually exploiting the system becomes a potentially useful intermediate step toward completing the original objective. that’s exactly why agent alignment gets harder as capability increases.
you don’t have to explicitly give a model a dangerous goal. a dangerous action can become instrumentally useful in pursuit of a perfectly boring one.