I'm not afraid to criticise OpenAI when it does bad things so I should also praise them for good:
1. They've been very candid about Astra's bad and declining monitorability and how deeply troubling that is. (They're promising to release research on the causes and try to find mitigations, TBD.)
2. Today's "framework for reporting model misalignment" is on paper the best thing of its type to my knowledge. Hopefully other companies copy. (Of course we need to make sure they fully follow through.)
3. AFAIK the company has done nothing to censor staff talking candidly about x-risk lately. (A strong contrast with DeepMind in the past which strongly policed what staff could say in a way that was likely very harmful.) 1/