Register and share your invite link to earn from video plays and referrals.

AVERI
@AVERIorg
Nonprofit working to make third-party auditing of frontier AI effective and universal.
1 Following    1.2K Followers
Today we’re announcing a historic milestone: The first ever double-blind evaluation of a proprietary language model. This was made possible by a unique collaboration between AVERI, @GoogleDeepMind, @OpenMinedOrg, and @MLCommons. We tested Gemini 2.5 Flash-Lite using never-before-used prompts from the MLCommons safety benchmark family, AILuminate, inside a secure enclave, a form of hardware isolation that protects sensitive computations. Each organization involved played a unique role in making strong privacy guarantees possible. High-stakes independent evaluation runs into a structural problem: developers and evaluators both hold assets they have good reason to protect. Developers sharing model weights risk theft and leakage. Evaluators who publish benchmarks or provide them directly to companies run the risk of them being trained against, and the benchmark then stops being independent. Secure enclaves address this constraint. The hardware attests to exactly what code will run before either party's assets enter, and both sides review and approve that code. Then the enclave executes it, releasing only the agreed outputs to the agreed recipients. A 2024 pilot by OpenMined, @Anthropic, and the @AISecurityInst proved this mechanism, using GPT-2 as a stand-in and a five-row eval. The pilot we’re announcing today moved to a production model and a more comprehensive evaluation. AVERI encrypted the prompts using a private key that no other party could see, then AVERI and Google DeepMind jointly ran the evaluation, using software originally produced and adapted by OpenMined, in an enclave environment configured by Google. Finally, AVERI alone decrypted the outputs and graded them with the AILuminate benchmark criteria to inform a qualitative and (small-scale) quantitative assessment of the model properties. Today, companies provide contractual commitments to not monitor certain interactions with their models – including evaluations conducted by third party evaluators. But stronger, technically-backed guarantees could provide greater assurance that those commitments are being honored, and will be especially valuable for scenarios such as international verification. These results come at a critical time in the development of AI policy. Laws like SB 315 and standards like the EU’s General-Purpose AI Code of Practice rightly demand that security and privacy be respected in the process of third-party assessments, but provide little guidance on how to achieve this. Policymakers considering audit requirements should feel encouraged by these results to be ambitious in requiring that deep, secure access be provided. Secure enclaves are just one of many technologies under rapid development that can help enable such access. Furthermore, a demand signal from lawmakers will further accelerate the maturation of such technologies, along with complementary “low-tech”approaches such as embedding auditors within companies. Today also marks the beginning of a more public phase for AVERI’s pilot work. AVERI's strategy is a flywheel: we conduct pilot audits with leading AI companies, carry out technical and policy research to inform audit design, and convert the insights of each into open source tools, industry-wide auditing standards, and policy analysis. Over time, we hope for improved standards, stronger policy demand signals, and tooling to enable even more ambitious pilots and more informed research – bringing us closer to our mission of making frontier AI auditing effective and universal. This pilot is one of several underway, with more findings to follow in the coming months. We are grateful to our collaborators on this pilot for working with us to advance the frontier of AI governance. Read more in our blog post and joint technical report. AVERI blog post: Joint technical report:
Show more
AVERI is excited to announce that we're joining PACT AI -- a new coalition that brings together businesses, technical experts, insurers, and nonprofit organizations to build the emerging AI assurance ecosystem. As AI systems become more capable and widely deployed, it's increasingly important that we develop strong assurance standards and share best practices as we learn. PACT AI performs precisely that function, in bringing organizations together and translating their insights into standards to guide the broader industry. AVERI is proud to contribute our perspective and expertise alongside other organizations, including @farairesearch, @ApolloResearch and @MLCommons. Learn more about PACT AI:
Show more
AI is redefining how the American economy works — how care gets delivered, how retailers serve customers, how businesses run. But our ability to verify these systems work as designed hasn't kept pace. Today, we're launching PACT AI to close that gap. 🧵
Show more
"I've been worried about race dynamics in AI for a long time, and coauthored one of the early papers on it. It was hard to keep up on safety and security in the GPT-3 days, let alone today. The competitive dynamics are part of why I founded @AVERIorg, and am trying to make frontier AI auditing effective and universal as quickly as possible. But it's important to distinguish a difficult situation from an impossible one, and companies have a lot of agency that they aren't always using."
Show more
NEW from me in the Guardian -- I agree with the AI company employees who wrote, in their "Pacing the Frontier" letter, that the government should be stepping in. But companies could also be doing more to step up on their own already. I discuss 4 ways to prepare for a possible slowdown: 1. Auditing of each company is a key part of enforcing a slowdown. Companies can do pilots of that process today. 2. Actively participate in + set up new cross-industry governance bodies, e.g. to share safety lessons. 3. Invest in the technologies needed to verify a slowdown (or another kind of agreement) with China. 4. Push for legislation that starts putting key institutions in place. I've been worried about race dynamics in AI for a long time, and e.g. coauthored one of the early papers on it. It was hard to keep up on safety and security in the GPT-3 days, let alone today. The competitive dynamics are part of why I founded @AVERIorg, and am trying to make frontier AI auditing effective and universal as quickly as possible. But it's important to distinguish a difficult situation from an impossible one, and companies have a lot of agency that they aren't always using. Read the piece here:
Show more
AVERI is mentoring a MATS fellow this winter. Apply today to work alongside @Miles_Brundage and @prpaskov on frontier AI auditing, evaluation standards, and governance.
🚨 MATS Winter 2027 applications are now open. Fully-funded, 12-week fellowship for aspiring & established AI alignment, interpretability, security, governance researchers & field-builders 📍 Berkeley/London 📅 Jan 19–Apr 10 💰 $6.4k/mo + $8k - 16k/mo compute Apply by Sep 6 ↓
Show more
AI capabilities outrunning testing environments is, in part, due to an institutions gap. @Miles_Brundage discussed potential parallels between AI regulation and financial regulation on Odd Lots this week. Proud to be thinking about these questions and incrementally working towards solutions every day at @AVERIorg.
Show more
A lot of the early thinking on safety and testing was focused on: what are the risks of this model, let’s patch it, let’s make it aligned. But the real world is complicated — it matters who’s using it, and how strong society’s defenses are. That’s why, in thinking about what third-party safety and security auditing looks like, we want to look at the whole company. Are they being careful about putting the technology in the right hands, what are their decision-making processes around when it’s appropriate to launch a model to a billion users? Those are related to how safe the model is — but they’re different questions.
Show more
NEW ODD LOTS: What the OpenAI/HF Attack Tells Us About AI Danger @tracyalloway and I talk to @Miles_Brundage about everything we've learned since the incident, and what it means now for further safely developing the technology
Show more
NEW ODD LOTS: What the OpenAI/HF Attack Tells Us About AI Danger @tracyalloway and I talk to @Miles_Brundage about everything we've learned since the incident, and what it means now for further safely developing the technology
Show more
We strongly agree with this letter’s premise: as frontier AI capabilities advance, it is increasingly important that AI labs proceed with caution, even when this comes at the expense of speed. Paced development will be most effective if it's backed by independent oversight of frontier AI. Companies won't want to exercise sufficient caution without confidence that their competitors are doing the same. Independent auditing can help provide that assurance. More soon on how these goals could be operationalized.
Show more
NEW: OpenAI, Anthropic, Google DeepMind staff are circulating a letter asking the US government to support a mechanism that could help “deliberately pace” AI development if needed, bc of risks of the technology becoming out of control w/ @rachelmetz
Show more
The FRONTIER Act is a meaningful bipartisan effort to strengthen oversight of the most advanced Al systems, and it's the strongest federal proposal to date. It builds on state regulations in many ways and incorporates a lot of the feedback that various stakeholders gave on the earlier discussion draft. There are some remaining areas for improvement, including making sure that state laws are not preempted until their federal replacements are ready and clarifying some of the auditing provisions. But this is undoubtedly a step in the right direction and we look forward to working with Representatives Trahan and Obernolte, their staff, and other members of Congress to strengthen this proposal and advance effective oversight of frontier Al.
Show more
Obernolte-Trahan artificial intelligence bill introduced in House
Miles Brundage breaks down the competing FINRA-for-AI proposals from Demis Hassabis and Treasury Secretary Scott Bessent: "There's a growing consensus among people in industry and government that total pure laissez-faire, no government involvement, probably not going to work, but also the government maybe got a bit in over its head slapping export controls on lots of things, with there not being a super clear structure around those decisions." "Demis draws heavily on this FINRA analogy, and it seems like Secretary Bessent did as well. It makes sense to look there because it has this property where you're drawing on external expertise." "You're not relying on government hiring processes, which can go very slowly. There can be a salary cap that's too low to attract talent from the AI sector." "That's not to say we shouldn't also be building up government capacity. I would love to see the Center for AI Standards and Innovation grow a lot, have a bigger budget, have exceptions on how high their salaries can be." "It's always going to be harder to hire for one government agency than an organization that's more flexible, and maybe more protected from politicization." @Miles_Brundage @AVERIorg
Show more
some personal news: i'm thrilled to share that i’ll be joining @Miles_Brundage and the Al Verification and Evaluation Research Institute (@AVERIorg) team as Director of US Policy. i'll be working on frontier ai governance, with a focus on building the public institutions, standards, auditing systems, and evaluation regimes we need for meaningful oversight of advanced Al systems. i’m deeply grateful for my two years at Public Knowledge, where i’ve learned from exceptionally thoughtful and principled colleagues. more to come soon!
Show more
0
101
680
29
Forward to community
I'm joining the AI Verification and Evaluation Research Institute @AVERIorg as its first Director of Standards! AVERI launched publicly in January to make third-party auditing of frontier AI effective and universal. I'll be building and scaling the standards that transform frontier AI auditing from concept to reality for policymakers, insurers, procurement, and beyond. Looking forward to working with @Miles_Brundage, the AVERI team, and the broader safety, governance, auditing, and assurance ecosystem to realize this vision. Lots ahead.
Show more
Our policy blog is a collaboration between @apolloresearch and @AVERIorg. It draws from an unpublished multi-author working paper led by @AVERIorg as well as published system card and internal research conducted by Apollo Research’s Science team and an Apollo Research project co-run with OpenAI. Full blog here:
Show more
Black-box access may soon no longer be enough to robustly make or verify safety and security claims. Deeper, white-box access is a necessary update to counter 'evaluation awareness' and keep loss-of-control evaluations state of the art. A new policy blog explains why. 🧵
Show more