Register and share your invite link to earn from video plays and referrals.

Joshua Saxe
@joshua_saxe
Now: cofounder @ Abundant Security. Before: AI | cybersecurity @ Meta. Earlier: social science, classical/jazz piano, hacking scene
1.2K Following    4.7K Followers
This OAI/HF piece from @GaryMarcus and Zack Korman -- whose main takeaway is that AI labs should practice defense in depth and suffer repercussions when they don't -- is representative of the response from many security folks. This position is not even wrong, it's just that it repeats the obviously true security catechism of the past 20 years in the face of a milestone event that marks the beginning of a new era in our field. It's as though we've discovered Winter is Coming and we're mainly focused on the banal politics of who's to blame in how we found out. To be clear: should OAI/Anthropic/Irregular/Meta/etc. do the things we've been saying forever and pay a price when they don't? Yes. Is it dangerous that they're not doing them? Yes. But also -- thought leadership should engage the new questions -- like: - What do we do about the fact that your median regional hospital network has *worse* security than OpenAI, Anthropic, Huggingface, and Meta? How do we secure tens of thousands of critical organizations quickly? What's AI's role here and what isn't AI's role here? - What are the potentially destabilizing geopolitical implications of hacking agents that can turn amateur-level non-state cyber actors into elite state-level actors now? - What will cybercrime look like when amateurish ransomware affiliates suddenly have the capabilities of elite nation states and what do we do about this? - To what extent will we need to guardrail our networks from internal loss of control incidents now given the business pressures to turn more and more engineering functions over to agents?
Show more
Finally listened to this; I liked the technical discussion, but disagreed pretty strongly with the societal and catastrophic risk analysis * @RyanGreenblatt gives a compelling intuition for why the next four years of AI progress could be as transformative as the last four. He also makes a convincing case that competitive pressures will lead organizations to deploy imperfectly aligned systems, producing real economic and moral harm through reward hacking and misalignment * But that does not by itself get us to an all-out civilizational catastrophe or a permanent loss of control. * To make that leap one has to explain why the market and policy feedback mechanisms that helped societies navigate previous technological and economic transformations e.g. oil spills, nuclear meltdowns, industrialization, enclosure, and so on, will fail in the case of AI. * The catastrophic argument seems to depend on two assumptions that are at least in tension. First, organizations must find AI systems safe, reliable, and valuable enough to place them in positions of extraordinary power - and must do so at a scale capable of creating civilization-level risk. Second, those same systems must be sufficiently overtly or covertly misaligned to cause a catastrophe. * If we see a steady drumbeat of incidents resembling the OAI HuggingFace episode, it seems unlikely that businesses and governments will continue handing these systems ever-greater control without stronger safety and security controls, or that labs wouldn't respond to this pressure (or to public / legislative AI safety advocacy) * One way to resolve this tension is to imagine that advanced AI will behave well enough to earn deep institutional trust and then suddenly defect once it has acquired sufficient power. Ryan seemed to offer something in the direction of this scenario explicitly. But I find it much less plausible than a world in which harms emerge more continuously, producing visible signals and eliciting market and regulatory responses along the way. * To his credit I found Ryan's takeover scenario at about the same level of gloss as the broader literature here from Scott Alexander, Zvi Mowshowitz, Yudkowsky, Soares, etc, none of which I've found convincing! And Ryan didn't put all his probability mass on this scenario occurring (Soares and Yudkowsky seem to) * But anyways, this leads to my broader meta-critique; AI safety is dominated by people with technical backgrounds, even though its central forecasting questions are substantially economic, political, and sociological. * We know AI is technically dangerous (as are nuclear weapons and many other technical objects in our civilization). The crux is what people will do with, to, and around AI and these are not technical questions. * How much risk will firms tolerate when deploying increasingly capable but unreliable systems? How will courts, legislatures, etc, respond, as evidence of harm accumulates? * In this vein, AI safety may be too much of a disciplinary monoculture to answer its own most important questions. * We need more economists, political scientists, historians of regulation, and sociologists in these conversations. When those perspectives are included, they often produce a substantially different, and generally less catastrophic, view of the future.
Show more
- It's urgent and indispensable that we fix AI cybersecurity policy now - I'm linking the slides from my keynote at the AI security forum below. I'm really passionate about the argument in the slides and I think the integrity of our social fabric depends on something like this happening soon. If you're reading this without knowing me until recently I was Meta's senior technical expert on AI security; I worked on frontier model evals, AI driven defense of Meta's infra, and was in lots of policy discussions with the other labs and the US/UK governments. From this I became deeply convicted that the frame policy folks are using to understand AI cyber risk is wrong from first principles; given where we're at in AI cyber capabilities this error will be very costly if we don't correct it. In the current frame, safety is imagined primarily as a property of individual models, and individual model launches are treated as the main objects of risk and the main opportunities for intervention (e.g. blocking a model launch). But the main object of risk is actually the softness of our entire national IT infrastructure (and therefore society) in the face of AI cyber capabilities, and the main object of intervention is to *increase the net benefits AI's dual use capabilities can offer to defenders while minimizing attacker uplift*. Government intervention should focus on using whatever methods -- AI or otherwise -- to mass inoculate society as fast as possible from the upcoming onslaught of cheap superintelligent hacking agents while also using these hacking agents to help do this and while minimizing their benefits to attackers. To be clear: I'm saying we should treat cyber risk as a public health concern and with an early-pandemic level of urgency. This would involve robust public infrastructure for surveilling the readiness of our economy and critical infra, understanding what's working for defenders and what's trending among attackers, and shaping policy at scale and with nuance around bending risk downwards. There are some efforts moving in this direction, notably those within CISA; but federal cyber defenders need to be empowered with more scope and scale and we need to go far beyond this. There are also important efforts inside the AI labs. But the labs can't impose regulations requiring, say, the boards and CEOs of tens of thousands of companies to properly fund the step change in cyber defense budgets that's needed right now. There really is an indispensable massive role for federal government and executive leadership here in keeping us all safe. After I gave my talk yesterday I did 1:1s with AI security folks from the labs, US CAISI / AISI, large AI safety grantmaking orgs, etc, who were in attendance. I think there's general agreement from our community here, and so a lot of this is up to our elected leaders, but perhaps there are ways the AI security community reading this can help catalyze this...
Show more