Sometimes models starts aligning w bad human behavior
“Higher levels of deception”
Saachi Jain, OpenAI’s head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas compared with its predecessor and wasn’t reliable enough to safely release. The model performed poorly on tests measuring alignment, or how well the model adheres to what humans would like it to do. Specifically, GPT-6.1 Astra showed higher levels of deception: It wasn’t always honest about telling users of the actions it did or didn’t take.
顯示更多