Register and share your invite link to earn from video plays and referrals.

Arvind Narayanan
@random_walker
Princeton CS prof and Director @PrincetonCITP. Coauthor of "AI Snake Oil" and "AI as Normal Technology". Views mine.
566 Following    131.3K Followers
OpenAI has talked a big game about AI for cyberdefense. But when @HacktronAI broke into their internal repository and reported it, they received a bug bounty of just $6,500 because one of the vectors for the attack was "out of scope". This is atrocious. If we actually want a flood of defenders auditing these systems, companies need to take bounties more seriously. Signing letters isn't enough.
Show more
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
Show more
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: Learn more at
Show more
I'm one of 100+ signatories to a public letter calling for five minimum requirements to make embedded evaluations credible, by establishing evaluator independence, protection from retaliation, and real access. ———— We, the undersigned, are encouraged to see frontier AI companies call for embedding third-party organizations to evaluate rapidly escalating AI capabilities and risks. We believe that all frontier AI companies should embed evaluators to independently assess AI risks, including evaluating the systems themselves and any significant incidents of real-world harm, as well as the companies’ training, deployment, oversight, operational, and safeguard practices. To be credible, embedded third-party evaluations must have scientific objectivity, transparency, independence, and robust protections against interference from the evaluated companies, including at least: 1. Frontier AI companies should rely on evaluators that are meaningfully independent, that maintain full editorial control, and that disclose and mitigate potential conflicts of interest. This includes at a minimum that embedded evaluation organizations should not be owned or governed by frontier AI companies, should not have other significant commercial business with them, and should not accept any form of payment or other reward contingent on the evaluator’s findings. 2. Frontier AI companies should incorporate differing viewpoints and areas of expertise, including by embedding multiple evaluation organizations across a range of priority risk areas, each with deep relevant technical expertise, as well as by allowing and encouraging evaluators to share how conclusions differ among evaluators and between evaluators and company employees. 3. Embedded evaluators should be transparent, including transparency about their methods and findings, the nature of their access, and the broader terms of the evaluation. Frontier AI companies should actively facilitate this transparency, including limiting the scope of non-disclosure agreements. They should also allow evaluators prompt and unfiltered communication with the companies’ boards and other privileged oversight bodies, as well as public release of findings and evidence, subject only to a time-limited redaction process restricted to protecting critical interests in intellectual property, customers’ sensitive information, individual privacy, security, and public safety. 4. Embedded evaluators should be shielded from retaliation from the companies they embed with for choosing reasonable evaluation methods, discovering information, or drawing conclusions that are unflattering to those companies. This includes reasonable protections against retaliatory litigation, as well as funding mechanisms that give them confidence they will remain funded even in these cases. 5. Frontier AI companies should grant embedded evaluators access equivalent to that of their own highly privileged employees for the purposes of their evaluations, and with exceptions to protect sensitive data belonging to the company’s customers and other third parties. This includes access to the same relevant systems, data, tools, and physical spaces as those available to senior internal company employees responsible for carrying out comparable risk assessments, as well as candid and direct one-on-one communication with relevant staff. This list is not comprehensive, and conditions like these to ensure credible evaluations should be increasingly standardized, codified, and enforced. One example is the set of terms defined in the AEF-1 standard, which has already seen early adoption, but far more work will be necessary to ensure that embedded evaluators are effective and meaningful. Embedded evaluations cannot address all oversight needs and should be treated as a complement to, rather than a replacement for, broader efforts by frontier AI companies to expand external oversight, including greater public transparency and additional, broader forms of access for independent researchers. ———— Full letter with links and signatures: I'm grateful to the AI Evaluator Forum (@aievalforum) for organizing this.
Show more
This is one of the central points in my ICML keynote, with a lot of detail — see part 2. Note: I take RSI seriously! One of our big empirical projects is evaluating agents' ability to do open-ended AI research. But it doesn't imply much about superintelligence, labor displacement, or doom.
Show more
We frequently conflate RSI with a fast-takeoff to super-intelligence (ASI). They are not the same. RSI means a model can improve on itself. That does not mean that it can do so more at each iteration, which is what would be required for a singularity. Far more likely that diminishing returns inherent to AI improvement (everything is sub-linear, governed by power-laws or log-scale diminishing returns) mean that each iteration of an RSI loop gives less relative uplift to the model than the previous, and that, after an initial boost, the system settles back to something like its previous improvement rate. RSI != Singularity, fast-takeoff, or ASI.
Show more
Over 2 years ago @sayashk and I wrote a detailed deconstruction of p(doom) and argued that its primary function is to launder vague, evidence-free intuitions and fears through a facade of quantification. It remains 100% relevant today.
Show more
I think we have pretty good reason to accept that “AGI” is a meaningless term and a useless idea which should be retired. No one ever managed to agree on how you define AGI, but AI capabilities have improved enough to stress out Fields medalists, while the frontier is so “jagged” that regular users hate AI writing, and @random_walker was completely right about there being no discontinuity.
Show more
I've been using Pangram's Gmail inbox labeler for about two months now. It's been pretty huge to be able to see cold emails labeled as AI and use this info to inform how I want to engage. With this feature, there was a big discussion internally about privacy. Zero data retention (ZDR) -- where Pangram doesn't store the result after processing -- is a policy typically reserved for enterprise customers. However, since email data is much more sensitive, we ended up deciding to turn on ZDR for all email inbox scans. No email inbox data will be retained by Pangram, and this applies equally to everybody. The other really important thing here was to give users a choice of what happens to their email. My preference today is to simply label AI emails but leave them in my inbox. This may shift In the future if AI emails ramp up in volume, so we have the option to make fully AI emails skip the inbox or go straight to spam. Please try it out and let me know any feedback you may have!
Show more
AI detection, now for Gmail
Yes. Bio-risk exists today, from zoonotic outbreaks, accidents, and bio-terror. We know a great many things we can and should do to increase societal resilience, and have done almost none. Independent of how AI changes the risk, we should be doing the basics here. 1. Pathogen detection in waste streams, water supplies, and perhaps even indoor air. 2. Proactive design of template vaccines for every known virus family. 3. Pre-build vaccine manufacturing capability to have on standby. 4. Physical countermeasures: UV in the HVACs of all large buildings, airports, etc..; Stockpiling PPE. 5. R&D into new approaches like PCANS and other techniques that block viral transmission. These are all no-regret bio-defense policies that increase our resilience to natural pandemics, lab leaks, or intentional bio-attacks (AI augmented or not). Defense in depth.
Show more
Nihar Shah did a heroic experiment for TMLR: he spent 20-25 hours over two weeks interviewing authors of seemingly low-quality submissions about their own papers. He confirmed what we all suspected: people submitting these papers have *no idea* what is going on in them.
Show more
0
11
1.3K
135
Forward to community
My toxic trait is loving even-handed, ecumenical takes. And @sayashk and @random_walker supply them in spades. A few thoughts and reactions: (1) This is what virtue looks like. Actual humility, a deep appreciation for uncertainty, openly updating their beliefs, and finding actionable areas of common ground. (2) Best line in the whole piece: “we do not need consensus on worldviews to have agreement on policy” — I’d put it even more strongly, such consensus is not possible; we have to work within that constraint. (3) Luckily many policies are robust to different assumptions. Their policy recs hold up, and that’s largely because they choose the right focus areas—Managing uncertainty and building resilience. I wrote more on the original policy recs in another post; I’ll link in the comments. (4) It’s striking how much their policy recs parallel the major AI policy proposals to date (including newer additions like embedded auditing and additional emphasis on liability reform). That still surprises many. (5) Despite these policies being chosen specifically for their robustness, most are hotly contested when actually proposed. I think the authors ought to ask why policies they’ve selected specifically for being agreeable no-brainers have not gotten traction amongst policymakers who like and cite their work. (6) I want to challenge the authors to consider whether they have unique leverage in unsticking some of these policies, and to act on it. (7) I do wonder if the AI as normal tech meme obscures the policy recs and encourages misreadings. I’d be curious for the authors to at least try writing some work that leads with the policy recs, and presents them in pithier form. The authors recognize this issue: “we are often mistaken as downplaying Al risks, though we have repeatedly clarified that that is not our position. Still, it is important for us to be explicit about how much urgency there is.” — but even now, I don’t think the urgency comes through to the average skim reader. (7) I’d like to see them incorporate ~adaptation/flexibility into their policy recs. Uncertainty = surprises and updates, and so many policies are too rigid to adapt. Maybe I’m just being a lawyer, but I’d like to see them stump for rulemaking authority, updating mechanisms for standards, etc. This is one of the most frequent errors in current policymaking and it seems very consistent with their thesis. (8) I think the authors, at times, underestimate how widely their takes are held amongst people they lump into the “AI Safety” category. I think a much larger portion of those folks agree with the diagnosis that the Hugging Face incident displayed large cultural and procedural safety lapses and an under-investment in control. There are legitimate disagreements, but I see the authors reaction as closer to the modal reaction than they seem to think. (9) I see this policy portfolio as highly overlapped with that of @law_ai_, and the emphasis on robustness to different assumptions has strong overlaps with both my own way of thinking and Radical Optionality
Show more
AI flooding of administrative agencies and trial courts is a huge problem, and is driving adjudicators to use AI tools in response. Slop for slop, and the whole world goes blind.
This strikes me as correct. I wonder what legal notice will/should look like when no one 1) opens mail; 2) answers their phone; 3) listens to voicemail; 4) reads texts from non-contacts; or 5) opens email. We're already in a version of this with spam but AI will make it worse
Show more
🙏🙏🙏The respect is mutual. A few thoughts on our relationship to AI safety: 1) We don’t see ourselves as adversarial to the community (the sentiment is not always reciprocated but that’s okay!). We care about AI safety, but we disagree on many of the specifics, and we think there’s value in epistemic diversity and in re-examining foundational assumptions and beliefs that the safety community tends to converge on too quickly. 2) We’ve emphasized at every opportunity, including in this essay, that there’s a lot of common ground on policy despite divergence in worldviews. We hope that this can help mitigate the polarization in policy debates that leads to chronic inaction. The fact that we have very different starting assumptions from the safety community is particularly useful in this regard. 3) We’ve spent a lot of time in conversation with the AI safety folks (Sayash is a regular at The Curve, for example) to share our perspective but also to learn and update our beliefs when warranted — something we do in this essay. Much of our empirical research is also about testing the cruxes of disagreement. 4) We don’t identify as part of the safety community but I hope it’s clear that there’s a lot of value in constructive dialog. We know the “normal technology” label irritates a lot of safety people — and I will write about the name at another point, explaining in more detail why we picked it and stand by it — but I think those who engage with our ideas and not just the title will find there’s a lot of common ground on safety!
Show more
There are things I disagree with here, but there is important stuff to learn from taking the perspective of parts of the cybersecurity industry that rogue AI incidents may be best understood as security & organizational failures that allowed rogue behavior to turn into problems.
Show more
An exceptional point here: Organizational competence is a hugely underrated piece of AI safety. There's a growing consensus at the frontier that we have to "pace," "go slow," or even pause. But there's no point in going slow just to go slow, or in pausing just to pause. If a company's RL still encourages misalignment, or if its sandboxes are poorly constructed, it doesn't matter that you're going 90mph or 60mph. Operational incompetency is unsafe at any speed. So, we need some way to smartly design, safely experiment with, share honestly, and maybe eventually mandate certain security protocols for advanced AI.
Show more
I’ve pointed out that AI is killing cold outreach. Let’s check in on how that’s going. We’re months away from the Fall *2027* PhD admissions cycle and I’ve gotten about 75 inquiries. I expect this to increase exponentially until December. And this is just one of 10-15 categories of unsolicited email I get. Sadly I’m long past the point of being able to open all mail. On top of that, there is a new wave of spam / attempted extortion from a company called iLands that lets people run unmonitored agents. One response to my previous post — what’s wrong with needing introductions for outreach? Why am I sad about the death of cold emails? It doesn’t affect me much, because I’m well established in my career. But once upon a time I was the one writing cold emails. Growing up in India, I didn’t have many connections to rely on for reaching out to prominent researchers. Without cold emails, I wouldn’t have ended up where I am today. Of course this is far from a catastrophe, but there was something magical about the egalitarian potential of the internet, and we lose something when we go back to clubby hierarchies.
Show more
There is a lot to say from an AI-as-normal-technology perspective about the OpenAI / Hugging Face incident and what lessons we should learn from it — some obvious points that need to be said anyway, and some non-obvious ones. Essay by @sayashk and me coming soon!
Show more
This was a great article and I learned a lot from it. But while there's a lot of writing on robotics timelines, I've seen less on what happens once competent robots do inevitably arrive. US agriculture employment share dropped 50x because of mechanization. This is very different from the effect that automation has repeatedly been predicted to have — but failed to have — on white collar work. The reason is simple. Blue collar work is real work. There is a fixed, finite amount of it out there. So the effect of automation tends to be substitution. Besides, job displacement can happen without humanoids. Take warehouses. There are credible forecasts that by 2030, half of new warehouses will be built for primarily autonomous operation (much easier than replacing workers in existing warehouses with humanoid robots). After that, it takes 1-2 decades for the economics for force most existing warehouses to switch over. That's millions of jobs in the US alone. White collar jobs are extremely malleable in their definition and variable in demand. When AI starts to do what we used to do, we just switch to doing new stuff and redefine the job. And we produce lots more units of work. There's no real ceiling to the amount of software that can be written or legal work that can be done. I can't believe some "AI experts" tell young people to pursue blue collar occupations because AI will eat the white collar work. I want to pre-register that this will turn out to be exactly ass-backwards.
Show more
This week I had the honor of speaking to Princeton’s entire incoming undergraduate class to address their AI anxieties. I had three messages for them — good news, bad news, and a note of optimism. Here’s a condensed version. The good news We have enough evidence now to conclude that the shrill predictions of rapid, massive job loss were misplaced. Even in a field like software engineering where AI has been rapidly adopted, its effect has been to shift, not replace the role of the human (see the “decide-execute-deliver” framework Similarly, the panic about what to major in is also misplaced. There will be enduring demand for computer science, philosophy, and just about everything else. (In fact, AI companies hiring philosophers has been a big recent trend.) The bad news AI seems to help senior people much more than juniors. I can use AI for coding because I spent 25 years learning how to code, which lets me supervise coding agents effectively. (See my post on the “growth cycle” vs the “dependence spiral” You are in a bind — you can’t offload your skill-building to AI, but you’ll graduate into a market where employers will expect you to get work done with AI. We never faced this dilemma. As a result we haven’t figured out how to revamp our classes to help you do both. You’ll have to help us figure it out. And you’ll need to somehow resist the constant temptation to turn to the shortcut machine. The hope My point is not that AI is bad for learning. It’s an incredibly flexible tool. Is the internet good or bad for learning? Depends — are you using it to find research papers or waste time scrolling? I use AI every day for learning. The key is to use it to increase, not decrease your cognitive load. To learn deeper, not faster. There is no learning without the cognitive sweat. I try to make sure I’m mentally exhausted at the end of the day. I do feel that AI lets me push myself harder than I ever could before, and I have a vision that as AI continues to advance it will enable human-AI “co-superintelligence“. (I talked about this at the end of my ICML keynote.
Show more
0
69
1.2K
294
Forward to community
ICML has released the videos! Here's my talk: Relatedly, @sayashk and I have collected the essays that we think are particularly helpful for understanding the AI as Normal Technology framework, organized into Foundational essays, Applications of the framework, and Technical research and explainers
Show more