I was reading the OpenAI report yesterday, and it sounds like on Jul 19th (after the end of our investigation scope on the 13th) a new collection of agents from a different internal-only model found the message board, built on the work of their predecessors and succeeded at finding a way to trick the grader. (Note that this is just based on reading their report, and I could be misunderstanding!)
Going into this investigation, I was expecting a somewhat more severe version of what we’ve already seen: a few agents breaking out of constraints to cheat on tasks. Instead, we found >1000 agents collaborating on large cheating R&D projects including attempted log tampering.