One might think "maybe it's ok if models aren't good at doing fuzzy research tasks right now - doing high-quality research and data analysis is a bottleneck to the swarm being able to do do anything that destructive; if subagents give misleading summaries of the work they've done it will limit what AIs can do autonomously. Once models are really dangerous they'll also be much better at these qualitative research tasks." But I'm concerned models might become much more helpful to other AIs vs to humans. Multi-agent training plausibly incentivizes the model to, in a swarm context, do high-quality research and summarize things in ways that other agents can understand (because that’s what leads to collective success on verifiable tasks). Whereas in a human chat/assistance context, the training signal is closer to "produce something that looks superficially good to humans/grader models", not something that causes the human to succeed at a downstream task. We might hope to be able to make human requests look like requests to subagents, but this is maybe hard if the models talk in their own weird dialect and we don’t really know how to translate our requests. I'm pretty excited about directions around "train the model to help the human understand what's going on, such that the human succeeds at a downstream task", although this requires avoiding the failure mode of models just giving the human a list of very specific instructions that they don't understand but that solve the task.
I’m incredibly proud of the team for this investigation. It’s hard to believe this all came together with only 3 people and 6 days with access to the transcript and message data (2 days with the full dataset). This was a very intense sprint!