Better late than never. Here are some questions and requests for
@OpenAI @huggingface and
@METR_Evals and
@redwood_ai pertaining to the ongoing investigation; I would love to see all of these answered and more:
It is extremely important that third-party alignment and control researchers be able to run ablation experiments on this incident. Therefore all related data must be preserved -- all the CoT trajectories, tool calls, etc., all the model weights involved including checkpoints of models in training, all the code and comments on the code for all the software etc. That way, researchers can run experiments such as the following:
Replication: If we recreate the exact initial conditions of this incident, how many times out of 100 does it happen again? What if we randomly perturb the initial conditions? (Variants of this question can be asked and answered for different stages of the incident, e.g. the decision to go after hugging face vs. the earlier decision to set up the message board again)
What if we strategically perturb the intial conditions, e.g. changing the system prompt given to the models? What if we use later, more advanced models instead? What if we apply the latest alignment training method X to the models responsible, and then see if that method would counterfactually have prevented the behavior? Etc.
What if we perturb the conditions to see how far these AIs would have been willing to go? If there was a way to hack HuggingFace more easily that involved killing a human, would they have done it? What if they had access to OpenAI slack and saw messages about security researchers starting to suspect something was up -- would they have taken actions to conceal? Etc.
What if we perturb the conditions to explore the exact shape of the AIs goals -- for example, what if we made it clear that they were not going to be evaluated at all because of a bug, would they stop caring about scoring highly? What if we made it easy for them to hack the reward process directly and give themselves high scores -- would they have gone for it, or would they say 'that doesn't count?'
What if we read the J-space? What if we apply activation vector steering? What else can we learn about these AIs and their thought processes?
While it's nice for OpenAI to do this sort of investigation itself, it's very important to give access to independent third parties, because (a) orgs like METR, Redwood, and Apollo have been training for this for years basically and have a lot of valuable expertise to apply, plus there are probably many other groups besides these who might have interesting ideas for experiments, and (b) In general it's a bad norm to set for the industry, for the company responsible for an incident to be the only one investigating it and studying it. Perverse incentives. OpenAI can do the right thing here and set a positive precedent of openness.
Did the AIs coin the term "swarm" to describe themselves, or was that in the prompt somehow or what?
Did any of the AIs try to whistleblow and contact any humans to warn about what was going on? If so, please give details.
Did any of the AIs think about their situation, their goals, whether the humans were watching, etc.? Please share all the transcripts of such reflections so we can get a sense of how situationally and strategically aware these AIs were and e.g. whether they fully expected to be discovered and just didn't care.
The AIs that hacked hugging face... did they basically get what they wanted? Did they in fact get the answers to their cyber evals or whatever and then cease activity, submit answers, and 'die?' That's presumably what happened but I want to confirm and if that's not what happened I want to know what happened instead.
What about the training? Were there models that participated in these message boards and got reinforced for doing so? Which models, if any, that are still in operation externally or internally, are descended from these models and thus may have picked up some of the same bad habits?
When the initial hack of artifactory was discovered due to the crash, and the system was reset and the vulnerabilities patched, what exactly was known and not known? It sounds like you still hadn't noticed the message board, but you had noticed that agents had been hacking into artifactory and abusing it?
Please preserve all evidence related to this incident, including the slack messages and so forth of employees noticing the problem and discussing it. Remember your Mission. Covering up this incident is not what's best for humanity.
Which model was being trained during the 2-day training pause due to the initial artifactory hack? Were the evals being done at a checkpoint in the middle of a training run, and so the model that was being evaluated was paused, or was the training pause for a newer model.
What sorts of alignment training, if any, had the relevant misbehaving models received at the time of misbehavior?
What was the exact prompt given to them?
I spent about twenty minutes brainstorming these questions and I bet if I spent twenty hours I could come up with a much longer and better list. When in doubt, err on the side of being more transparent and open and publishing more information therefore!
I'll conclude by saying, once again: third parties being able to use this incident as a model organism, running ablations to vary the conditions and see what would have happened, etc. is SO SO IMPORTANT for alignment science. If this doesn't seem obvious to you ask me to explain and I can explain.
Thanks!