1) How can we make a safety case regime paired with embedded evaluators effective?
2) The AI industry aspiring to a standard of safety cases is good; notably the term “safety case” as used by other industries implies a high level of rigor. AI cos/safety orgs have only published safety case sketches thus far.
3) I disagree with Dario’s essay where he says he thinks pacing measures like compute are more gameable measures; I think a company allocating 80% of its compute to make critical infrastructure resilient to cyber attacks or curing cancer would be easily verifiable and hard to game, as well as good for the world.
Safety case regime is *the* thing to aim for but the science of risk assessment behind them is still being developed. So we shouldn’t only rely on the safety case regime (this will take a long time to get sufficiently right); we should also use compute (or other RSI inputs) measures.
Today, OpenAI is launching a new Alignment Research blog: a space for publishing more of our work on alignment and safety more frequently, and for a technical audience.