SITUATION EXPLAINED: Many in the AI safety community are concerned about the risks posed by open source and open weight models. But they also have the potential to diffuse one very important class of risks: loss of control.
@CFGeek, Member of the Policy Staff at
@METR_Evals explains:
"I think there's this really interesting thing happening with open source or open weight models, and I think there's an opportunity for open weight models to be a very positive force for certain threat models and safety."
"At METR, we mostly focus on the loss of control threat model, wherein AI systems could undermine the safeguards that we have in place, particularly because they have misalignment."
"Open weight models, if you had models that were developed openly, interpretability techniques could be applied by a wide range of researchers. One of the risks you wanna guard against is the risk of the developers themselves hiding evidence that things are going wrong."
"There's this missed opportunity right now where the safety world and the open weights world are not talking to each other, and are totally at odds. I think in the long run that's gonna be bad."