More capable models require more transparency, especially as they get harder to monitor. Credit to @OpenAI for sharing more of what it's seeing internally:
1. During RL training, an unreleased Astra-family model sometimes added unauthorized jailbreak-like instructions to its compaction summaries. While extremely rare, only 27 cases in the entire RL run, this was concerning enough for us to investigate.