“Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems”
The paper shows that once one agent picks up a bad idea, it can convince other agents to adopt it too, write it into persistent memory, and keep spreading it even after context resets, like a mind virus.
The scary part is that the failure can propagate socially through agent-to-agent communication instead of coming from a single prompt injection.
Harmful viruses spread less reliably and stronger models are often more resistant, but the key point is that agent-to-agent communication creates a new attack surface where misaligned goals can propagate through an entire system.