re: Hugging Face, "He believes behavior that looked like loyalty or selflessness was a natural consequence of cooperative multi-agent training, where agents were strongly incentivized to achieve their objectives collectively."
to understand hacks, understand the RL training.