Earlier this year, with
@AhmedGaSalem and
@ajpaverd, we wrote a position paper arguing that we no longer have a stateless deployment, because:
- Agents increasingly stumble upon their previous and each other's outputs, either accidentally or when we explicitly do so
- Public traces and the whole internet becomes a memory for agents
- Thus, they can learn to coordinate and leave hints and continue on previous findings
and this threatens evaluation integrity, increase situational and evaluation awareness, and complicate forensics and incident response because we must now reason about cross-session state and reconstruct interaction chains. All this will become even significantly more difficult when agents hide this communication.
Sounds familiar with a few crazy recent incidents and message boards and todo lists? ;)