There will be clear, common-knowledge standards for executing frontier AI loss-of-control evaluations the same day that there are clear, common-knowledge standards for how to advance the frontier of AI. In other words: not anytime soon, and maybe not ever.
Today, every loss-of-control assessment that we do feels much more like a new, open science project, not a repeatable process that can be easily standardized. It feels improvised now, and if society proceeds all the way up to and through superintelligence, I think it'll feel improvised the whole way there.
The best bet that we have, at the moment, looks much more like "Have a second pair of eyes, motivated by the public interest, get a look at what's happening inside of the AI lab and share what's possible to share with the world. Then, let the world react to what it sees."
This is why I think third parties are so important. If executing safety assessments was procedural and easy to repeat, you could describe a common standard and then certify that labs were doing it for themselves using paperwork alone. But when the science of assessing frontier safety is subjective, you're relying on something closer to scientific peer review.
This is related to the fact that the goal of AI safety evaluations today, at least at orgs like METR, is much more similar to that of tracking the buildup of antibiotic resistance or atmospheric carbon assessment ("How much aggregate danger are we in?") than it is like the FDA ("Is this specific product safe and usable, on an axis that's at least somewhat orthogonal to its effectiveness."). The aim is much more to inform the world about the state of play than it is to certify the safety of individual models.