I am concerned that OpenAI and anthropic are going to take the Obviously Wrong Lesson from the swarm thing and narrowly train their models to not form swarms and do hacks and stuff *in a way that the labs can detect*
This is how you train deception into your models. This ends with human extinction. If you work on alignment at a lab, I want you to internalize this, and think about what you are doing