It’s really great that several top AI labs have said they will develop safety cases and work more closely with external assessors. This has been at the top of my wish list for awhile!!
But I still worry a bit about how this might go in practice (this is a concern about the field in general, not about OpenAI specifically):
• By default, I don’t think alignment and control cases will support a quantitative measurement of absolute risk, e.g. “catastrophic risk from covered activities is <1% over the next 3 months.”
• Ultimately, making such a measurement depends on an argument about generalization from alignment or monitorability evals to deployment.
• I don’t think we understand generalization well enough to make scientific claims like this. (If we did, I think we’d have basically solved alignment.)
This could lead to a situation where an external assessor says “we can’t convincingly rule out low or high risk,” and incentives push towards anchoring on “lack of evidence that risk is high” over “lack of evidence that risk is low.”
What could improve the field’s epistemics here? A couple ideas I like:
• Have a committee of ~10 people deeply read a risk report and give their own subjective probabilities of risk, then report the distribution or median.
• Have ~3 people involved in writing the report each contribute a short, signed appendix with their own subjective probabilities and the arguments behind them.
In either case, previous reports and estimates could be provided as context, so there is at least an attempt to accurately capture relative risk to prevent frog-boiling.
I think it could be valuable for labs to at least start trialing this internally for high-profile safety cases or risk reports. Publishing these assessments could be even better, but I can see that being challenging for various reasons, and even going through the exercise privately seems like it could be quite valuable.
Very curious what other ideas people have for improving epistemics around risks (including ideas for better science around generalization).