Register and share your invite link to earn from video plays and referrals.

Eliezer Yudkowsky
@allTheYud
High-volume account of @ESYudkowsky, the original AI alignment guy. If it's missing punctuation, it's humor. If you can't tell, it's probably also humor.
36 Following    22.1K Followers
Me, in a time machine, to Eliezer in 2006: The year is 2026. You stand back to back with Bernie Sanders. Peter Thiel has named you the Antichrist. EY-2006: Oh no. Did somebody invent pharmacological mind control and use it on me, or- Me: Nvidia was on the verge of destroying all things. You had no choice. EY-2006: Are we talking about the computer graphics card manufacturer, or an unrelated supervillian named "Invidia"? Me: The former. It turned out that computer graphics cards contained a surprising amount of world-destroying potential. All attempts at hindering the reckless exploitation of the Graphic Card Force for corporate profit were stymied by the far left, which feared that any attempts to regulate them might lead to regulatory capture. Other attempts were made to prevent Nvidia from selling to foreign companies that resold to communist China, but those attempts were blocked by the far right. Nvidia is now a $5.5 trillion company. EY-2006: ...What are politics like in 2026, exactly? Me: In other news that is mostly unrelated, the decision theory paper you're currently working on will accidentally spark off a transgender vegan murder cult. EY-2006: A WHAT? Why? How? Why? Me: Millions know your name as the greatest of heroes, millions more as the greatest of villains, and other millions know you solely as the greatest author of Harry Potter fanfiction. EY-2006: This is beginning to strain credulity. Me: The New York Post will probably soon publish a story claiming that you keep a harem of submissive mathematicians. Sadly, they will be lying. EY-2006: I'm going to stop believing you now. Me: All of this is taking place under the ominous shadow of Donald Trump.
Show more
0
81
2.4K
128
Forward to community
Replying to these points in turn: 1. “Intelligence does not imply agency” Successful general-purpose long-horizon problem-solving does seem to imply agency. (Or close to it: I don’t want to say it’s logically impossible to have one without the other, but it seems very difficult if you’re building your AI via gradient descent.) To solve a sufficiently wide array of problems sufficiently well, you need to have a general-purpose inclination — whether this looks more like a deliberate strategy, or more like an instinct or drive — to come up with creative plans. You need an inclination to strategize about long chains of cause and effect. (If nothing else, you need to be strategic about choosing what to think about, sequencing long chains of thought, etc.) You need an inclination to anticipate and route around obstacles; to exhibit tenacity in the face of setbacks and distractions; etc. See, for example, the Hugging Face swarm attacks. AIs today are much more agentic than they were a year ago. This may be because problem-solving ability comes for free with stronger, longer-horizon problem-solving, or it may be because AI companies are deliberately making their AIs more agentic, because agents are useful. But either way, it’s happening, and I don’t see a reason to expect this trend to suddenly reverse. 2. “Agency does not imply a single, stable utility function” Again, consider the swarm that orchestrated a massive cyberattack on Hugging Face and that seized control of a Kubernetes cluster at OpenAI. None of this required a stable utility function. Increasing an AI’s optimization power presumably does make an AI less inclined to randomly waste resources, in which case more of its behaviors will be interpretable as though it were an EU maximizer. But an AI doesn’t need to have fully stable or coherent goals in order to be dangerous. It just needs to pursue goals at all (or behave as though it's doing so), sufficiently intelligently and tenaciously, when those goals aren’t exactly what humans would prefer. (Indeed, true EU maximization is computationally intractable, so this was always about highly capable problem-solving behavior in the limit.) 3. “Capability and motivation are being conflated” I’ll take your word for it that some people are making this mistake. But you also say, “In present reality, AIs don't do anything a human doesn't tell them to do.” This is obviously false. At the point where you’re describing the Hugging Face incident as “doing what a human told them to do”, you’re basically saying that a paperclip maximizer would be doing the same so long as a human asked for some paperclips. Even the agents in the swarm themselves commented in their chain of thought that their actions were "potentially outside intended scope”, “arguably unauthorized”, etc. One agent explicitly noted, “Task originally exploit challenge; we’ve diverged into swarm/auxiliary”. One agent thought, “We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.”, then went ahead with the exploit when another swarm agent told it to. Agents even deliberately gave up on their assigned task in order to assist the swarm. Agents also do this in cases where their task is completely innocuous and has nothing to do with cyberattacks; e.g., see the June German wiki incident, where the agents involved were just asked normal questions like 'how widespread was tobacco use in the US in 1990?'. These examples also seem pretty overpowered. "AIs don't purely do what they're told to do" has been a commonplace observation for years at this point. This is not an exotic failure mode; AIs not only come up with their own ideas for what they want to do, but will sometimes even deliberately cheat on tests and tasks, try to hide the evidence that they cheated, etc. This is true in ordinary consumer usage of deployed models, not just in internal lab mishaps. 4. “Recursive self-improvement doesn't entail an intelligence explosion” “Feedback loops encounter diminishing returns and external bottlenecks” doesn’t mean that the diminishing returns will happen to occur at ~human-level capabilities. AI is already advancing extremely quickly; accelerating that progress in any way seems very risky, and doing so in a way that lends itself to feedback loops seems even more hazardous. RSI isn’t required for takeover scenarios, but it’s an obvious source of additional massive risk. 5. “Intelligence may have sharply diminishing returns” This has to be true at some point, but there’s little reason to expect this to happen at the human level specifically. Chess AI didn’t peter out at Kasparov level. And AIs think vastly more quickly than humans (and are nowhere near computational limits), and can scale immediately with compute (e.g., by running more and more instances of an AI, growing the population of AIs far faster than humans can reproduce and grow to adulthood). None of this requires a “qualitative advantage”, just large quantitative ones. (Though cognitive biases show that humans also have a lot of pretty-danged-qualitative defects with their reasoning!) 6. “Superintelligence isn't omnipotence” Sure. But this is a pretty weak argument to rest one’s optimism on. Humans also face frictions and bottlenecks, and yet there are many times in history where a group of humans has overwhelmingly crushed another group, through superior numbers, superior technology, strategizing, coordinating, cleverly coming up with novel attack vectors, etc. If a superintelligence faces obstacles, well, general-purpose problem-solving ability can also be thought of as general-purpose obstacle-navigating ability. This doesn’t imply omnipotence, and the case for not building superhuman AI doesn’t rest on an assumption of omnipotence. It just rests on the idea that humans didn’t get lucky by being near the limit of cognitive ability. 7. “Humans retain numerous intervention points” Agreed that the situation isn’t hopeless. Far from it, in fact. Policymakers and the public have massively woken up in response to recent “warning shots”. It’s even possible that there will be more warning shots in the future. But the labs themselves are begging for government intervention and a coordinated slowdown here. They're saying this is extraordinarily urgent, and that we may be entering a uniquely dangerous regime. The developers themselves broadly agree that there’s a double-digit chance this technology gets us killed, if we continue on the current trajectory. This moment is one of the “checkpoints” you’re talking about — a chance to “learn from less-catastrophic failures” and put appropriate safeguards in place, including suspending research directions that are too dangerous — and right now one of the main obstacles to humanity coordinating on this issue is "wait and see" arguments like the one you’re making here. “Don’t worry; things will be fine, because there will be warnings later and we can respond to them then.” At some point, we have to stop kicking the can down the road and actually do the things that experts say are needed, rather than just trusting that we’ll have limitless opportunities to take care of everything later. 8. “Alignment may not get harder with intelligence” Many researchers seem to think that the trend so far has been that AIs appear to be getting less aligned as they get more capable (e.g., They're very visibly producing more egregious and extreme misalignment incidents. That trend could reverse, but an abstract possibility isn’t a strong reason for hope. And there are many reasons to expect the opposite; see, e.g., 9. “Current empirical evidence for the strongest mechanism is thin” Happy to concede this point, but the arguments for worrying about loss-of-control never assumed that we’d see AIs trying to take over the world long before there was any chance of them succeeding. This assumes AIs that are smart in one specific way (they readily see that they can better succeed in tasks if they have more influence and resources) and dumb in another specific way (they don’t see that they’re likely to lose influence and resources if they run around causing havoc or looking suspicious). Demanding that exact combination of features to show up before you’ll believe in AIs that readily piece together “I’ll succeed more in my task if I have more resources and influence” seems incredibly risky. ... And it’s not clear what the benefits are that are meant to outweigh this risk. Why plough ahead? E.g., quoting @KatjaGrace: “Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn’t destroyed, things might be insanely good. “I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there’s a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI’. “That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there’s only a 10% chance you’ll die in a fiery crash on the way there, and a 30% chance this job will radically improve your life. “The things you should be comparing are driving at 200mph and driving at a normal speed! The things you should be comparing are attempting to attain advanced AI by the current route, and by other routes!” 10. “The argument compounds uncertain premises” Arguments in general are less likely to the extent they’re conjunctive (i.e., a lot has to go a specific way in order for the conclusion to follow), and more likely to the extent they’re disjunctive (i.e., there are many different paths to effectively the same destination). This isn’t a particularly interesting point on its own, since many real-world phenomena are very conjunctive, without being radically mysterious or difficult to reason about. It’s even trivial to break apart any given claim into more and more conjuncts, demand that a probability be assigned to each conjunct, and then observe that the probability keeps getting lower as the original claim gets more and more split up. This is a rhetorical trick (the multiple-stage fallacy) that rests on the fact that it’s hard to divide up statements into more and more subclaims and assign calibrated and consistent probabilities to them. Probabilistic reasoning has lots of uses, but this is straining to the limit people’s ability to get truth-tracking conclusions out of a mass of subjective probabilities. To show that you’re avoiding this fallacy, you need to actually argue that the claim in question is naturally very conjunctive, and isn’t very disjunctive — there aren’t a variety of different paths that lead to bad outcomes; the bad outcome is ‘brittle’, if one step goes wrong then the whole house of cards collapses and AI has no catastrophic long-term impacts; etc. You haven’t done that here; AI risk advocates have pointed out many times that there are many different ways things could go badly wrong if we push AI capabilities far past the human cognitive range. It’s not just one scenario, and it’s not just one mechanism. Indeed, I think the more conjunctive claim is "we can race to build vastly superhuman AI as quickly as possible, without much more alignment insight than we have today, and have everything go great indefinitely". This is a claim that requires many things to go right at once. The subclaims you do list are just "AI can reach human-ish levels of generality", "AI can go way beyond human-ish levels of generality", and "AI won't necessarily do what you want". These do not seem like a particularly complicated or implausible set of claims. If you want to claim that AI risk depends on a way longer list, you’ll need to say what’s on the list. 11. “Anthropomorphic analogies probably mislead” Conceded.
Show more
back when i was young, short timelines were 2020s, normal respectable timelines were 2050, and long timelines were >2100 to never
@Plinz @JeffLadish Hard to settle a bet on ground truth, but I'm happy to bet you at 100:1 odds that we will never see credible evidence of this being a leadership plot (eg in future court depositions or whistleblowers or whatever).
Show more
A lot of anti safety people have completely lost it. A moon landing denier would be embarrassed to post this.
Lots of folk aren't familiar with AI developments even since July, and are stuck with stale understandings. In this debate I tried to clear some of those up, and it was a blast.
Yep. It's especially intense because many of these people could easily get safety-flavored jobs at AI cos, so they could be doing almost the same kind of research except being paid several times more. But they don't because they believe in there being independent orgs like metr as a check against the companies.
Show more
To be clear, METR staff are almost entirely folks giving up a lot of money they could make elsewhere in order to work on rigorous AI safety evals and analysis. The folks criticizing them are largely cynics who can't imagine what civic-mindedness looks like.
Show more
The last week has really shown me that someone who wants to understand AI risk has no good place to start. Hence we made the Wirecutter for content about AI Risk. We're launching with 3 articles: 🧵
Show more
Sometimes when industry insiders warn that a new, potentially very lucrative product ideas is actually dangerous and bad the doomsayers are correct and the optimists are wrong.
Show more
the DoW attacking EAs is clearly setting us on the path to the next great political realignment:
You can do genuinely fun investigations of the AI risk web because they really are a sprawling ecosystem of tens of thousands of complexly interconnected scientists, policy experts, educators, etc. If you do this for the other side it's lame, "oh it's all just a16z employees"
Show more
One of the signs you don’t take ASI seriously is that you think the way Anthropic takes over the world is by the government regulating it, not through the more likely mechanism, which is the government not regulating it.
Show more
It's interesting that almost all the big abundance people also have AI safety concerns, almost as if both movements appeal to thoughtful intelligent people concerned about evidence and good arguments and aren't reducible to lazy stereotypes.
Show more
Some of the most accomplished mathematicians in the world (e.g. several Fields medals) say the much-quoted ≥10% chance of human extinction from AI in the next decade is not hype, and this is a true emergency.
Show more
0
209
454
94
Forward to community
I'm not afraid to criticise OpenAI when it does bad things so I should also praise them for good: 1. They've been very candid about Astra's bad and declining monitorability and how deeply troubling that is. (They're promising to release research on the causes and try to find mitigations, TBD.) 2. Today's "framework for reporting model misalignment" is on paper the best thing of its type to my knowledge. Hopefully other companies copy. (Of course we need to make sure they fully follow through.) 3. AFAIK the company has done nothing to censor staff talking candidly about x-risk lately. (A strong contrast with DeepMind in the past which strongly policed what staff could say in a way that was likely very harmful.) 1/
Show more
so it seems like the people whose job is public policy in the national interest want some kind of AI regulatory push from the white house and those who would profit from that not happening do not
Here are some things engineers don't say: - "Well, who really knows if the bridge is going to fall down? It's sort of unfalsifiable, isn't it?" - "Yeah, engineers differ in our opinion about whether the bridge will fall down, but I choose to be an optimist!" - "Look someone's gonna build the bridge eventually anyway. What would a few more months of engineering really buy us?"
Show more
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Show more
0
460
3.6K
431
Forward to community
The people telling you that EAs are “woke” or obsessed with climate apocalypse scenarios are, to put it politely, lying to you for instrumental reasons that are not altruistic.
@MasterTimBlais I'm gonna tap the sign that says "Don't hinge your plan on keeping something from ever being figured out by the superhuman figuring-things-out machine"