100% agree, would also love to see details on how the monitoring system didn't catch this directly, did the model kinda hide the fact that it was hacking hf?
OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field learn from. Did the top-level agent know about the hacking, or was there some "value drift" between it and its subagents? How did it rationalize its behavior?