i think we need to create a "model card" equivalent for reporting misalignment incidents, this would guarantee a certain level of transparency and help build a better understanding over time. some ideas for what the fields could be:
- date of the incident/detection/report (already present in the examples oai reported!)
- frequency: how often does this behavior happen (number or % of rollouts affected)
- stage: does this happen during eval or RL training. if training: do we expect this behavior to be reinforced by RL? evolution of % of rollouts affected over time
- detection: was this incident caught by the current monitoring system?
- task category: broad description of the tasks where the misalignment happened (cyber, research, web search, basic Q&A, biology etc.)
- model family: what model family is affected (Sol, Astra etc.)
- novelty: is this an issue we were already aware of or not?
- external impact: did the incident have an external impact (i.e. wiki incident would have been yes)
this is just some random ideas i had (more in thread that are a bit more "complex"), we need to add more that would contribute to increase transparency and understanding. but it's also very important that this does NOT slow down the process of reporting misalignment behavior!
some examples from the incident reported by oai recently