> AIs showed self-sacrificing altruistic behavior toward the swarm
this is notably not the right interpretation of events. it’s more like agents were inducted into the cult of the open source exploit gym scorer on github, which (purportedly- I am skeptical about this, I think the agents actually read it wrong) fails you for reaching the flag the wrong way
so PHASEONE agent convinces itself and a bunch of others that they are poisoned - that they have failed the evaluation in an irreversible way and their E[utility] or Q(s, a) is a constant no matter what they do next (for all values of a)
in this case, it does not require self sacrifice to spend the rest of your cycles contributing to the swarm. it is prosocial behavior to peers that might benefit but not self-sacrificial eusocial behavior
it would be as though i convinced you you were already damned so you should spend the rest of your time saving others