SITUATION EXPLAINED: Cybersecurity concerns have figured prominently in recent high profile AI policy decisions, including Anthropic's decision to hold back Mythos and the Trump Administration's decision to place an export restriction on Fable 5, now due to be lifted tomorrow. How confident should we be in the accuracy of the evaluations that are being used to make these decisions?
We asked
@CFGeek, policy staff at
@METR_Evals
"Big picture, AI risk assessment as a whole is in triage mode. There are many more risks and risk pathways than there are evaluators to do it, or unsaturated benchmarks to measure what's going on. And the capability development is just going apace."
"Cyber is one domain you see this in. You see this in bio. Saturation of the traditional question answering benchmarks, and you have to rely on other ways of measuring things like uplift evals, which just take a bunch of time."
"How are you gonna decide whether to release the next model based on putting a bunch of undergrads in a wet lab and having them do uplifting things for three months? It's just not feasible on a deployment type timeline."
"All of the traditional benchmarks are saturated or heavily caveated in some way. But ultimately you have to make a decision about whether to deploy the model or whether it's the right time to start coordinating with other developers on some sort of mutual safeguard regime."