Register and share your invite link to earn from video plays and referrals.

Chris Painter
@ChrisPainterYup
President @METR_Evals
1.4K Following    13.6K Followers
They left me out because a white dude with glasses had too much risk of Watney
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Show more
0
460
3.6K
431
Forward to community
METR is hiring in cyberforensics. We now embed researchers inside of frontier AI labs to stress test monitoring systems, assess AI loss-of-control risks, and investigate misalignment incidents. If you've investigated serious security incidents end-to-end, or managed teams that do, and want to apply DFIR skills in frontier AI, apply (and feel free to DM me with questions). Comp range is $400k - 580k cash.
Show more
There will be clear, common-knowledge standards for executing frontier AI loss-of-control evaluations the same day that there are clear, common-knowledge standards for how to advance the frontier of AI. In other words: not anytime soon, and maybe not ever. Today, every loss-of-control assessment that we do feels much more like a new, open science project, not a repeatable process that can be easily standardized. It feels improvised now, and if society proceeds all the way up to and through superintelligence, I think it'll feel improvised the whole way there. The best bet that we have, at the moment, looks much more like "Have a second pair of eyes, motivated by the public interest, get a look at what's happening inside of the AI lab and share what's possible to share with the world. Then, let the world react to what it sees." This is why I think third parties are so important. If executing safety assessments was procedural and easy to repeat, you could describe a common standard and then certify that labs were doing it for themselves using paperwork alone. But when the science of assessing frontier safety is subjective, you're relying on something closer to scientific peer review. This is related to the fact that the goal of AI safety evaluations today, at least at orgs like METR, is much more similar to that of tracking the buildup of antibiotic resistance or atmospheric carbon assessment ("How much aggregate danger are we in?") than it is like the FDA ("Is this specific product safe and usable, on an axis that's at least somewhat orthogonal to its effectiveness."). The aim is much more to inform the world about the state of play than it is to certify the safety of individual models.
Show more
This third-party incident investigation took enormous work not just from METR and Redwood staff, but also from people inside of OpenAI who enable third-party work. OpenAI shared over a thousand unredacted transcripts, and provided unusually high rate limits to let this team quickly analyze this extremely large volume of data. I think this exercise sets an excellent precedent for independent third-party investigation of misalignment incidents.
Show more
I agree + have long felt that the world outside of the Bay Area knows less about "AI control" as a concept than it should. Can you incentivize and assure good behavior from AI systems even if you know they are somewhat misaligned, by having them monitor each other?
Show more
Very proud of the work that our team has put in to fundraise the money needed for us to pursue ambitious assessment work while also maintaining a very high standard for funding independence. Very grateful to all of the people at METR and outside of METR who made this possible.
Show more
I had a great time on Hard Fork with @kevinroose and @CaseyNewton to talk about AI alignment. Today the stakes for AI alignment failures might feel small because we think of AI systems like junior virtual employees. Soon, when AIs are very capable and running large parts of our world in ways that we aren't closely observing, the stakes for alignment failures will be massive. As long as AI capabilities progress rapidly we may have very limited time to solve this problem or find ourselves between a rock and a hard place, where evidence of misalignment is pervasive, and yet we're stuck in a competitive race that requires making more and more advanced models.
Show more
I think people in tech, including employees, are still too weird about, deferential to, awestruck by, CEOs and co-founders of important companies. I think you’ll get a better picture of reality by modeling them not as “great people of history”, but more like “Your smart but fallible acquaintance who is very busy, was a bit prescient/‘early’ to their area, and suddenly making big decisions they deep down feel they’re kind of unprepared for but clearly the only person in a position to make”
Show more