Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators.
These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad.
We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes.
Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground.
To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should:
1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest
2. Rely on multiple evaluators with differing viewpoints and areas of expertise
3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings
4. Shield evaluators from retaliation
5. Grant access equivalent to that of highly privileged employees
There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum.
Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem.
See the public letter here:
Learn more at
The AI Evaluator Forum (AEF) welcomes recent statements regarding the importance of embedded independent experts to verify AI safety and security claims. This is a first step towards trustworthy oversight of frontier AI.
No single evaluator can do this work alone. A vibrant ecosystem of independent evaluators from distinct backgrounds can offer a range of expertise and methodology, help to ensure rigor, and avoid a single point of failure. The Forum exists to strengthen this ecosystem.
Realizing the benefits of third-party evaluations also requires independence, access, and transparency. AEF brings evaluators together to address these questions. Our first standard, AEF-1: Minimum Operating Conditions for Independent Third-Party AI Evaluations, envisions a minimum floor for access, managing conflicts of interest, funding relationships, recusal requirements, and transparency of evaluation terms.
Read more here:
The long-term success of third-party evaluation relies on standards and frameworks. AEF is committed to developing these with its members and collaborators. We welcome engagement from organizations committed to rigorous, independent work.
Learn more about AEF, express interest in joining, and raise questions for us here: