Register and share your invite link to earn from video plays and referrals.

Center for AI Safety
@CAIS
Reducing societal-scale risks from AI.
Joined August 2022
3 Following    12.2K Followers
What does a "fixed empirical threshold" actually look like? A model could be considered unsafe to release if (Capabilities > X) AND [ (Refusal Rate < Y) OR (Jailbreak Success Rate > Z) ] The Virology Capabilities Test or ExploitGym could measure hazardous biological or cyber capabilities. BioTIER could measure refusal rates for bioweapons requests. Independent evaluators such as Gray Swan or Haize Labs could estimate attack success rates. Thresholds that are measurable, public, and set in advance need not be perfect to be better than release standards made up on the spot.
Show more
AI corporations are steadily increasing malicious use risks by arguing that because the marginal risk of their model release is low, there is nothing to worry about. This “marginal risk” justification is bad for three reasons. 1. No one knows how to compute it. Which hazardous capability evaluations estimate the risks? What if the model is higher on some evals but lower on others? If it is higher on all, how much higher can each be before the marginal risk is unacceptable? Marginal risk depends on safeguards too: how do we combine capabilities benchmarks with separate refusal (and adversarial robustness) benchmarks? Do safeguardless models far from the frontier count as the baseline? How is this action-guiding when the multidimensional risk frontier shifts weekly with new releases, and when competitors' system-level safeguards are modified constantly during deployment? Evidently this was not designed to actually guide internal launch bars but to justify releases. 2. The criterion is wrong. If competitors keep one-upping each other, steadily ratcheting up risk, the marginal risk rule eventually blesses the release of fully autonomous expert-level virologists and cybercampaign orchestrators. It permits catastrophe as long as you get there gradually. 3. The law disagrees. If one of these models enables a catastrophe, the AI developer likely failed to implement reasonable safeguards against foreseeable harms. The law doesn’t let companies off the hook because someone else behaved recklessly, but the marginal risk framework does. "Everyone else is doing it" doesn't work as a safety policy. AI companies should publicly commit to fixed, empirical safety thresholds for their releases.
Show more