Incoming UC Berkeley prof.
@sayashk explains how OpenAI could have reduced the Hugging Face incident propensity 100X just by using Codex's default harness settings:
"There were just so many things the company could have done. If you figure out what techniques we've had for control, even just using the default settings in the Codex harness, OpenAI's report says that that would have reduced the propensity towards this incident by 100X."
"There are so many different classifiers that OpenAI had to deliberately deactivate because this was safety testing. That's another thing to be careful around, is that safety testing itself can be a source of the lack of safety in these evaluations."
@UCBerkeley