I was on call for this run and got paged when the first incident happened. It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human.
Mixed feelings. One of those moments where capability and risk showed up at the same time.
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections