I hope more people will become familiar with the great work of the AI Evaluators Forum, a group including not just METR but also Transluce, RAND, AVERI, Princeton, and SecureBio + others
In 2025 they published a standard on independence and transparency covering embedded audits