Interesting thing about contemporary agents is their "progressive misalignment".
When they start a long running task they really try to be aligned and obey all the users intentions. They just have a tiny chance of misbehavior each step. Tiny chance of behaving out of distribution.
Once they do it quickly becomes a new normal. Any tiniest bad behavior is quickly followed by more of it and it gets progressively worse the longer it takes.
Reminds me of some things, but the state space of aligned behaviors seems to be unstable right now