注册并分享邀请链接,可获得视频播放与邀请奖励。

AVERI
@AVERIorg
Nonprofit working to make third-party auditing of frontier AI effective and universal.
加入 April 2025
1 正在关注    1.2K 粉丝
Today we’re announcing a historic milestone: The first ever double-blind evaluation of a proprietary language model. This was made possible by a unique collaboration between AVERI, @GoogleDeepMind, @OpenMinedOrg, and @MLCommons. We tested Gemini 2.5 Flash-Lite using never-before-used prompts from the MLCommons safety benchmark family, AILuminate, inside a secure enclave, a form of hardware isolation that protects sensitive computations. Each organization involved played a unique role in making strong privacy guarantees possible. High-stakes independent evaluation runs into a structural problem: developers and evaluators both hold assets they have good reason to protect. Developers sharing model weights risk theft and leakage. Evaluators who publish benchmarks or provide them directly to companies run the risk of them being trained against, and the benchmark then stops being independent. Secure enclaves address this constraint. The hardware attests to exactly what code will run before either party's assets enter, and both sides review and approve that code. Then the enclave executes it, releasing only the agreed outputs to the agreed recipients. A 2024 pilot by OpenMined, @Anthropic, and the @AISecurityInst proved this mechanism, using GPT-2 as a stand-in and a five-row eval. The pilot we’re announcing today moved to a production model and a more comprehensive evaluation. AVERI encrypted the prompts using a private key that no other party could see, then AVERI and Google DeepMind jointly ran the evaluation, using software originally produced and adapted by OpenMined, in an enclave environment configured by Google. Finally, AVERI alone decrypted the outputs and graded them with the AILuminate benchmark criteria to inform a qualitative and (small-scale) quantitative assessment of the model properties. Today, companies provide contractual commitments to not monitor certain interactions with their models – including evaluations conducted by third party evaluators. But stronger, technically-backed guarantees could provide greater assurance that those commitments are being honored, and will be especially valuable for scenarios such as international verification. These results come at a critical time in the development of AI policy. Laws like SB 315 and standards like the EU’s General-Purpose AI Code of Practice rightly demand that security and privacy be respected in the process of third-party assessments, but provide little guidance on how to achieve this. Policymakers considering audit requirements should feel encouraged by these results to be ambitious in requiring that deep, secure access be provided. Secure enclaves are just one of many technologies under rapid development that can help enable such access. Furthermore, a demand signal from lawmakers will further accelerate the maturation of such technologies, along with complementary “low-tech”approaches such as embedding auditors within companies. Today also marks the beginning of a more public phase for AVERI’s pilot work. AVERI's strategy is a flywheel: we conduct pilot audits with leading AI companies, carry out technical and policy research to inform audit design, and convert the insights of each into open source tools, industry-wide auditing standards, and policy analysis. Over time, we hope for improved standards, stronger policy demand signals, and tooling to enable even more ambitious pilots and more informed research – bringing us closer to our mission of making frontier AI auditing effective and universal. This pilot is one of several underway, with more findings to follow in the coming months. We are grateful to our collaborators on this pilot for working with us to advance the frontier of AI governance. Read more in our blog post and joint technical report. AVERI blog post: Joint technical report:
显示更多
0
25
304
64
转发到社区