Register and share your invite link to earn from video plays and referrals.

Sayash Kapoor
@sayashk
Incoming prof @UCBerkeley AI agents, policy, evals, AI for science Essay/newsletter (AI as Normal Technology): Book (AI Snake Oil):
2.5K Following    15.2K Followers
POV on catastrophic cyber risk * We're in a new era, the right side of the cyber incident damages distribution will get fatter and elongated, making unprecedentedly large cyber incidents possible. * We'll continue to see the old regime of human bottlenecked cyber attack workflows too; this is *not* catastrophic and will account for most damages near-term. * But we'll now progressively uncover new categories of rare high impact risks for as long as offense improves and defense fails to price in the new risks. * What'll characterize the new fat tail is breadth in attackers' ability to leverage parallelized agent swarms (precursor: oai/hf), depth (hacktron's /goal based vuln research and exploit dev to hack openai, meta, slack), and motive (actors that want to do catastrophic damage, like geopolitical actors and nihilists; e.g. a Russian worm devastating the Ukrainian economy). * As @random_walker and @sayashk point out in their most recent piece, a fattening and elongating damages tail doesn't require superintelligence, only that capability keeps rising, and AI cost keeps falling, and defenders keep failing to price in the new risks. What we should do: * Treat each sample from the fat tail as a precious natural experiment and mandate disclosure. AI providers and victims withhold incident details; regulation should change that. OpenAI should be mandated to publish details of the Hugging Face hack as should all who got nonzero felonybench scores! * Use AI to finally make deployable defenses the industry has proposed for years and never implemented due to costs deployable: this includes current poorly implemented basics, and also as yet too-costly defenses like deception, allowlisting, and better runtime and binary program level protection. * Build sensors and actuators across inference providers, cloud GPUs, on-prem infrastructure and endpoints to interrupt agent inference by attackers and kill swarms and worms, catalyzed by federal mandated and international agreements. * Treat critical-infrastructure risk as an externality: breaches impose costs on society, which justifies regulation and subsidies. To be clear: I'm not arguing policy should center on catastrophe or even that we'll immediately see catastrophes. I'm arguing for a loop to address tail risks that merit concern: study each tail incident, feed lessons into strategy, and use automation to cut the cost of defense to prepare. Longer version:
Show more
OpenAI has talked a big game about AI for cyberdefense. But when @HacktronAI broke into their internal repository and reported it, they received a bug bounty of just $6,500 because one of the vectors for the attack was "out of scope". This is atrocious. If we actually want a flood of defenders auditing these systems, companies need to take bounties more seriously. Signing letters isn't enough.
Show more
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
Show more
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: Learn more at
Show more
My toxic trait is loving even-handed, ecumenical takes. And @sayashk and @random_walker supply them in spades. A few thoughts and reactions: (1) This is what virtue looks like. Actual humility, a deep appreciation for uncertainty, openly updating their beliefs, and finding actionable areas of common ground. (2) Best line in the whole piece: “we do not need consensus on worldviews to have agreement on policy” — I’d put it even more strongly, such consensus is not possible; we have to work within that constraint. (3) Luckily many policies are robust to different assumptions. Their policy recs hold up, and that’s largely because they choose the right focus areas—Managing uncertainty and building resilience. I wrote more on the original policy recs in another post; I’ll link in the comments. (4) It’s striking how much their policy recs parallel the major AI policy proposals to date (including newer additions like embedded auditing and additional emphasis on liability reform). That still surprises many. (5) Despite these policies being chosen specifically for their robustness, most are hotly contested when actually proposed. I think the authors ought to ask why policies they’ve selected specifically for being agreeable no-brainers have not gotten traction amongst policymakers who like and cite their work. (6) I want to challenge the authors to consider whether they have unique leverage in unsticking some of these policies, and to act on it. (7) I do wonder if the AI as normal tech meme obscures the policy recs and encourages misreadings. I’d be curious for the authors to at least try writing some work that leads with the policy recs, and presents them in pithier form. The authors recognize this issue: “we are often mistaken as downplaying Al risks, though we have repeatedly clarified that that is not our position. Still, it is important for us to be explicit about how much urgency there is.” — but even now, I don’t think the urgency comes through to the average skim reader. (7) I’d like to see them incorporate ~adaptation/flexibility into their policy recs. Uncertainty = surprises and updates, and so many policies are too rigid to adapt. Maybe I’m just being a lawyer, but I’d like to see them stump for rulemaking authority, updating mechanisms for standards, etc. This is one of the most frequent errors in current policymaking and it seems very consistent with their thesis. (8) I think the authors, at times, underestimate how widely their takes are held amongst people they lump into the “AI Safety” category. I think a much larger portion of those folks agree with the diagnosis that the Hugging Face incident displayed large cultural and procedural safety lapses and an under-investment in control. There are legitimate disagreements, but I see the authors reaction as closer to the modal reaction than they seem to think. (9) I see this policy portfolio as highly overlapped with that of @law_ai_, and the emphasis on robustness to different assumptions has strong overlaps with both my own way of thinking and Radical Optionality
Show more
The AI as Normal Technology guys are consistently some of the best and most interesting critics of a lot of arguments in AI safety world. I'm making a point to read everything they put out.
Show more
🙏🙏🙏The respect is mutual. A few thoughts on our relationship to AI safety: 1) We don’t see ourselves as adversarial to the community (the sentiment is not always reciprocated but that’s okay!). We care about AI safety, but we disagree on many of the specifics, and we think there’s value in epistemic diversity and in re-examining foundational assumptions and beliefs that the safety community tends to converge on too quickly. 2) We’ve emphasized at every opportunity, including in this essay, that there’s a lot of common ground on policy despite divergence in worldviews. We hope that this can help mitigate the polarization in policy debates that leads to chronic inaction. The fact that we have very different starting assumptions from the safety community is particularly useful in this regard. 3) We’ve spent a lot of time in conversation with the AI safety folks (Sayash is a regular at The Curve, for example) to share our perspective but also to learn and update our beliefs when warranted — something we do in this essay. Much of our empirical research is also about testing the cruxes of disagreement. 4) We don’t identify as part of the safety community but I hope it’s clear that there’s a lot of value in constructive dialog. We know the “normal technology” label irritates a lot of safety people — and I will write about the name at another point, explaining in more detail why we picked it and stand by it — but I think those who engage with our ideas and not just the title will find there’s a lot of common ground on safety!
Show more
There are things I disagree with here, but there is important stuff to learn from taking the perspective of parts of the cybersecurity industry that rogue AI incidents may be best understood as security & organizational failures that allowed rogue behavior to turn into problems.
Show more
An exceptional point here: Organizational competence is a hugely underrated piece of AI safety. There's a growing consensus at the frontier that we have to "pace," "go slow," or even pause. But there's no point in going slow just to go slow, or in pausing just to pause. If a company's RL still encourages misalignment, or if its sandboxes are poorly constructed, it doesn't matter that you're going 90mph or 60mph. Operational incompetency is unsafe at any speed. So, we need some way to smartly design, safely experiment with, share honestly, and maybe eventually mandate certain security protocols for advanced AI.
Show more
The biggest takeaway for me: “we do not need consensus on worldviews to have agreement on policy.” Whether you’re in the AI safety is a conspiracy camp, the AI is a normal technology camp, or the AI 2027 camp, there are many policy interventions we can implement today that would be good no matter what happens with AI. See, e.g., @IFP’s or
Show more
@random_walker Link to the essay: Really grateful to @joshua_saxe @snewmanpv @curl_justin @steverab @RodMoshtagi for feedback that dramatically improved the essay.
What does it mean to pace the frontier? Over the last month, @random_walker and I have analyzed the loss-of-control incidents at AI companies to understand what technical and policy interventions can improve safety and what companies should do to pace the frontier. The result is a new 13,000 word essay — our most substantial writing on AI safety since AI as Normal Technology. A summary of our arguments: 1) The polarization between the cybersecurity and AI safety communities is counterproductive. The safety community largely sees these incidents as a crisis for alignment, and worries that these incidents will become more damaging as agents become more capable. Cybersecurity practitioners largely see companies failing to take basic security precautions. We offer a middle ground between these communities as a way forward for improving AI safety. 2) We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents. But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions. 3) We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI. 4) We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community. 5) Organization governance should be a key tool for pacing the frontier. Unfortunately, AI companies are trying to reinvent basic aspects of organizational governance as a problem to be solved by improving the technology. But even developing better control techniques will not be enough if irresponsible individuals or teams within large organizations can choose not to use them. When a single misconfigured RL environment or unmonitored evaluation can cause real-world harm, individual teams should not be able to run potentially dangerous experiments without oversight from legal, security, and other teams. AI companies need processes for reviewing experiments, assigning responsibility for monitoring them, and investigating warning signs deeply before restarting experiments. If putting these processes in place requires pausing some experiments, companies should do so. 6) How should we reason about AI's impact on cybersecurity? It's plausible that advances in agent capabilities upset the offense-defense balance for cybersecurity. We cannot yet be certain, but there is enough evidence that agent capabilities might soon make widespread cyberoffense possible that urgent action is warranted. We discuss potential interventions for tilting the offense-defense balance towards defenders. 7) How our views have evolved over the last year. We take stock of AI progress and share how we have updated our views. In the essay, we did not pay sufficient attention to safety risks that arise during development and evaluation (as opposed to the widespread deployment of models). We were too confident that companies would take basic control precautions and underplayed the importance of jaggedness, which led us to underestimate how quickly capabilities could improve in domains such as cybersecurity. 8) At the same time, many distinctive claims of AI as Normal Technology have held up. In particular, we think recent incidents support our continuity hypothesis — the behavior of "rogue" agents became apparent and widely publicized while they are still incompetent at causing serious harm or hiding their traces. The societal reaction to even the relatively small harms from these incidents has been fierce (and the safety community deserves credit for keeping up pressure on companies). Whether this translates into meaningful changes in companies’ behavior remains an open question, and a test of the usefulness of the AINT framework. 9) In short, we’ve tried to synthesize the AI safety and cybersecurity communities' views into a coherent plan of action: hold companies responsible, invest in control, and strengthen defenses against specific risks.
Show more
0
45
344
106
Forward to community
The AI Evaluator Forum (AEF) welcomes recent statements regarding the importance of embedded independent experts to verify AI safety and security claims. This is a first step towards trustworthy oversight of frontier AI. No single evaluator can do this work alone. A vibrant ecosystem of independent evaluators from distinct backgrounds can offer a range of expertise and methodology, help to ensure rigor, and avoid a single point of failure. The Forum exists to strengthen this ecosystem. Realizing the benefits of third-party evaluations also requires independence, access, and transparency. AEF brings evaluators together to address these questions. Our first standard, AEF-1: Minimum Operating Conditions for Independent Third-Party AI Evaluations, envisions a minimum floor for access, managing conflicts of interest, funding relationships, recusal requirements, and transparency of evaluation terms. Read more here: The long-term success of third-party evaluation relies on standards and frameworks. AEF is committed to developing these with its members and collaborators. We welcome engagement from organizations committed to rigorous, independent work. Learn more about AEF, express interest in joining, and raise questions for us here:
Show more
There is a lot to say from an AI-as-normal-technology perspective about the OpenAI / Hugging Face incident and what lessons we should learn from it — some obvious points that need to be said anyway, and some non-obvious ones. Essay by @sayashk and me coming soon!
Show more
Until AI companies have adopted actual preventative measures, I'm afraid we are prone to many more Hugging Face incidents. Two big things that still seem missing: 1) Do not allow AI to make **any** code-merges into your control system unless they've been affirmatively cleared as safe. - If you don't pre-clear these code changes, they can *turn off* your other controls. This can make any of your after-the-fact responses, like paging, totally ineffective. - You should be using a high-quality monitor with CoT access for this. It is not enough even if you have a human review every pull request; you want to check the transcripts for misalignment evidence, and they're way too long. 2) Make sure your agent automatically fails-closed in very serious cases. i.e., implement a circuit breaker for if your monitor is getting blitzed by a bunch of attacks. - Otherwise, the attacker gets a very high number of retries, and can attempt very harmful things, so long as it is confident in sneaking _one_ through before humans can respond. We describe more inside Principle Four of Guidelight's Control standard: "Stop the AI from taking harmful actions even if it tried", but these are IMO the two most important practices. I will have a hard time believing the industry has taken the recent incidents seriously, until these measures are in place.
Show more
Understandably, the reward-hacking behaviors on display in the Hugging Face incident are being treated as a cyber-risk story. But I'm not seeing sufficient discussion of the fact that they're also an economic story. Reward-hacking behaviors pose a serious challenge for complex, long-horizon tasks in enterprise, considering that 1) by their nature, they are often unpredictable; 2) they are "fractal" in the sense that, even well-specified intermediate outcomes meant to guard against reward-hacking can themselves be reward-hacked (and so on at smaller scales); 3) they pose serious cyber, legal, and financial risks, as Hugging Face makes abundantly clear. As a result, you could say that reward-hacking is as bearish on the economic front as it is "bullish" on the risk front. It is greatly underappreciated by safety folks that, on average, the riskiness / unpredictability of a technology is going to be anti-correlated with its diffusion into the economy. This matters a lot in the AI case because capabilities growth is highly sensitive to the resources available for training, R&D and so on, resources that are at risk of drying up if AI produces insufficient returns, or only produces sufficient ones on the wrong timescale. This is connected to a form of fallacious reasoning that I often see from the AI maximalist camp, which is to assume that the requisite capital for AI training, R&D, and so on will always be available in arbitrarily large amounts, which is why it can be safely assumed that capabilities will continue to grow without fail. But this isn't true! The availability of capital for capabilities research and training is endogenous, and obviously so, to AI's economic prospects, even over very short timescales (one wonders, for example, how capabilities progress so far has been influenced by the timing and size of frontier labs' capital raises, relative to various counterfactual scenarios). I see far too little discussion among safety folks of how these two domains interact. If you're going to model capabilities growth, you also need to model the capital markets behavior that feeds into that growth. IMO reward-hacking is a productive site for thinking about this interaction.
Show more
AI capability vs. reliability is an extremely important issue. @sayashk makes a critical point by measuring reliability with different benchmarks (ex. agent accuracy rates) and finds reliability improvement greatly lags capabilities. AI labs should focus more on reliability now.
Show more
The reliability / capability gap is one of the most underdiscussed problems in AI, and yet is hugely consequential for everything from the future of work to x-risk. See my essay from June, which draws heavily on @sayashk's work :
Show more
Incoming UC Berkeley prof. @sayashk says AI companies have spent billions on alignment while investing orders of magnitude less in agent control: "If you think about interventions on control, I would argue there have been orders of magnitude less investments." "Companies have spent billions of dollars on collecting training data for their alignment efforts and aligning the models themselves." "But on things like ensuring the security of their sandboxes or making sure that their transcript analysis is done correctly or making sure that agents don't go unmonitored and take actions in the real world, I think that has been far less of a priority. At least before this incident, the industry's main point of intervention was focused on alignment." @UCBerkeley
Show more
Incoming UC Berkeley prof. @sayashk reveals AI reliability is improving 4–10X slower than capability and why 90% accuracy can still be catastrophic: "We came up with a list of metrics that other industries had used, industries like aviation and nuclear safety and so on. We tried to port these metrics to the AI agent space." "Reliability improving much more slowly than capability. Based on these early results, it was between four and 10 times as slow as capability improvements have been in terms of accuracy numbers." "That's why AI's effect on the job market has been augmentation not automation. You can't accept an error rate of even 1% or 5%." "Rabbit R1 and Humane Tech Pin were personal agent assistants. If your product orders your DoorDash food to the correct address 90% of the times, it's a catastrophic product failure because 10% of the times your users are mad at you." @UCBerkeley
Show more
Incoming UC Berkeley prof. @sayashk explains how OpenAI could have reduced the Hugging Face incident propensity 100X just by using Codex's default harness settings: "There were just so many things the company could have done. If you figure out what techniques we've had for control, even just using the default settings in the Codex harness, OpenAI's report says that that would have reduced the propensity towards this incident by 100X." "There are so many different classifiers that OpenAI had to deliberately deactivate because this was safety testing. That's another thing to be careful around, is that safety testing itself can be a source of the lack of safety in these evaluations." @UCBerkeley
Show more