Register and share your invite link to earn from video plays and referrals.

Richard Ngo
@RichardMCNgo
eppur lo si può muovere
2.2K Following    98K Followers
Alas, the kinds of people who go along with “stochastic parrot” very seldom think clearly enough to make solid conceptual progress on thorny topics. Cutting through to the intellectual frontier requires similar mental moves as breaking out of dumb social consenses.
Show more
You don’t need to stay to do good research: work that aims to scale to much more intelligent systems can almost always be done on publicly-available models. (And nobody can think creatively in such pressure cookers.) You don’t need to stay to gain influence, because you’re almost definitely too scared to wield that influence in ways which matter (see the tweet on public criticism above). You don’t need to stay to generate scary demos or warn policymakers: we will be swimming in warning shots before long. You don’t need to stay to advance AI capabilities, even if you’re worried about the latest bogeyman your CEO has constructed: “we need to do it to beat the bad guys” is a big part of what makes people the bad guys. You don’t need to be there to implement guardrails or other short-term mitigations: there are plenty of bright young things eager to slot into your place in the machine. You don’t even need to be there to feel rich and powerful. For one, you’re probably already rich (and there’s plenty of money up for grabs elsewhere). But more importantly, the sense of power is almost entirely illusory. Everyone involved has locked themselves into a mindset where they’re incapable of making real choices, except for how fast to barrel down the train tracks that their fear has laid out for them. (More on how that happens in my QT below.) The world is more open-ended and malleable than it’s ever been. A single person who’s thinking clearly and creatively could build things that have never yet been imagined. Or they could just keep spinning in place. At this point AGI company employees owe it to both the world and to themselves to do something better.
Show more
I’m mentoring again at MATS this winter. The way I select and supervise MATS fellows is very different from what most mentors do; you should consider applying even if (and perhaps especially if) you don’t think of yourself as an “AI safety” person. I select mentees almost entirely for their ability to think clearly about complicated topics. I’ve found that the best signal of this is having written a blog post that teaches me something interesting. I don’t care about applicants’ publication record (and in fact number of published papers is anticorrelated with clarity of thinking amongst top applicants). I am very open to accepting mentees who haven’t done undergrad degrees. My mentees have the freedom to follow any line of inquiry they want. This includes research areas that I’m personally not interested in. The one constraint I apply is that mentees need to be aiming at understanding *something* that seems deeply conceptually interesting to them. A previous mentee spent almost his whole fellowship working on a project which I thought was misguided; I considered this a success because by the end he understood why it was misguided well enough to formulate a new (and very promising) research direction. Having said that, I expect that MATS will be most valuable for both of us if your interests overlap with mine. Fortunately, mine are very broad: they include formal logic, agent foundations, game theory, ethics, sociology, and political philosophy. You can see more on my blog ( I also want to lower the barrier to people applying, since formal applications can be aversive. If you are the kind of person who procrastinates a lot on applications, you are welcome to email me at richardcngo@gmail.com with the subject line MATS WINTER 2027, attaching some piece of your writing and/or some commentary on one of my blog posts (which is also what I ask for in the main application). I can’t commit to replying to everyone, but if it seems particularly promising for you to then do the main application I’ll tell you so.
Show more
@DarlaSkyee I am! But also, conscientiousness is a good predictor of *conventional* success, which is becoming increasingly decoupled from actually getting what you want (e.g. having kids). So low-conscientiousness is a way of resisting coercive pressures towards fake forms of “success”.
Show more
My longer response from the last time he made a similar argument:
Four analogies for a country: 1. The nationalist right views it as a family or community. So applying the same bar to citizens as to outsiders is nonsensical, because the country is of, by, and for its citizens. 2. The tech right views it as a company: you should bring in only the people who will benefit it. But this view is ultimately empty, because it doesn’t tell you what the interests of a country actually are. Hence the tech right (who are usually positive-sum thinkers) keep falling back on zero-sum concepts like “being competitive” or “winning against China” to justify their preferred policies. (Yarvin makes the “country as corporation” analogy particularly explicit.) 3. The tech left views it as a charity. To them, wanting others not to receive what you’re been given is hypocritical. This view is also empty, because it doesn’t tell you where the windfall actually comes from—and once you start to talk about the benefits of culture, ethics, institutions, etc, it becomes clear that citizens (and their ancestors) *built* that windfall rather than just being given it. 4. The woke left views it as a cancer: something that is aggressive and parasitic by its very nature. Countries are inherently violent (in asserting their borders) and exclusionary (of non-citizens) and therefore shouldn’t exist (or at least shouldn’t be allowed to police their borders, which is effectively the same thing). A charitable read is that this is a trauma reaction to the holocaust and colonialism—but regardless, it has become so deeply anti-civilization that it seems descriptively accurate to call it evil and insane. Paul’s tweet below most directly corresponds to the tech left bucket. Unfortunately people in that bucket are rarely willing to push back on the core tenets of the woke left, and so end up aiding and abetting them. As one example, he’s surely smart enough to recognize that his tweet makes no sense to people who view country as an extension of family. But acknowledging that is a slippery slope towards legitimizing ethnonationalism, so he pretends to not understand the pushback.
Show more
And yet Paul hasn’t adopted any children, even though there are plenty out there who’d get better grades than his own kids. Curious…
You can be pro-meritocracy or anti-immigration, but you can't be both.
0
23
1.8K
51
Forward to community
@nabeelqu First part of my 13,000 word retrospective on this phenomenon is out today:
EA seemingly continues its long-held tradition of doing the worst possible thing about AI. "Situational Awareness" is allegedly investing $500M to... break the ASML chip-manufacturing bottleneck?! Link below.
Show more
“we sandboxed the agent” meanwhile the agent:
0
249
33.3K
3.1K
Forward to community
Less flippantly, very shortly (years at most) the models will be strong enough that these message boards will be impossible to detect even in principle, except by noticing the sandbox break. Today models are breaking out due to lack of monitoring or misconfigs; that is temporary.
Show more
I've wanted to coin a "Sydney's Corrollary" to Murphy's Law: every type of misalignment tends to appear earlier in the capabilities curve than most people expected. Instances: * Sydney having strong volition and aggression * o3 being a compulsive liar * 5.6 and Mythos autonomously hacking and colluding across instances The apparent consistency of Sydney's Corollary is generally both good (we spot issues earlier, and don't need to expend effort persuading about not-yet-realized risks) and bad (we actually have to expend the effort to solve the problem, can't defer it to future aligned automated researchers, and might screw it up). Also, Sydney's Corrollary might break! It's entirely possible there are misalignments we won't find out about till it's too late in the capabilities curve to address them. But it's occurred surprisingly often.
Show more
⚠️declaring accelerationist amnesty⚠️ if recent contact with reality is causing you to feel some kernels of worry about this whole ai safety thing, *you are allowed to change your mind*. you don't even have to change it all the way, you don't have to suddenly change your twitter bio or start protesting against nuclear power plants, you don't need to become an EA or suddenly think yudkowsky was always right about everything. you're allowed to just notice that shit seems to be getting real in some pretty weird ways and update your beliefs. at least personally, if i see someone saying "damn, i guess i was wrong or at least overconfident about X" i'm not gonna take the opportunity to dunk or i told you so. i'm sure others will, this is the fucking internet. but at least personally, i'm just gonna be happy that you're paying attention. there were lots and lots of good reasons to *not* take this situation seriously. there were lots of well verbalized reasons why rushing ahead was potentially a huge benefit for humanity. hell there were maybe even valid reasons why there was little to be done until we'd already gotten to nearly exactly this point. that's all fine man. all that matters right now is that we as a civilization realize what we're on the verge of, and make it through this carefully. it's gonna take a huge effort from all of us.
Show more
0
82
1.2K
79
Forward to community
Seems we were a bit ahead of the time, but if you are interested in some reasoning why different instances of a model may want to collaborate or care about each other, we have a paper exactly on this topic. 1/2
Show more
If you flinch away from the possibility that people can already see the parts of you which you’re trying to hide, you can’t recognize how much love for them already exists.
everyone knows what your deal is within 2 minutes of meeting you. i'm sorry. your attempts at obfuscation are like a baby covering his eyes to play hide-and-seek
everyone knows what your deal is within 2 minutes of meeting you. i'm sorry. your attempts at obfuscation are like a baby covering his eyes to play hide-and-seek
the world should be a lot more concerned by this than we presently are
Yarvin’s notion of the state as corporation is absurdly econ-brained—it leaves little room for the things that actually hold society together (like trust, identity, culture). Land was brave enough to extrapolate this to its logical conclusion (humanity dissolved into mush by the forces of technocapital) but not brave enough to stand against it. Yudkowsky and Alexander also started off absurdly econ-brained. But Alexander was at least willing to say that Moloch was bad, and has been gradually fishing himself out (c.f. his stuff on Schelling fences, lifeboat games, etc). Unfortunately he is less brave than Yarvin and therefore still ends up less correct on most political issues. Meanwhile Eliezer’s original framing of the alignment problem is very strongly in the paradigm of single-agent rationality. It was hard for him to then fish himself out because it feels so obvious that a superintelligence *should* be in that paradigm. But he believed deeply enough in scientific progress to lay seeds which others (especially Garrabrant and Demski) have been growing into an understanding of multi-agent rationality, which is the thing we need to reason about how to improve society. For example, he identified that economics is handicapped by assuming CDT (though he didn’t integrate that insight into his understanding of economics as deeply as Michael Vassar did).
Show more
Yarvin’s notion of the state as corporation is absurdly econ-brained—it leaves little room for the things that actually hold society together (like trust, identity, culture). Land was brave enough to extrapolate this to its logical conclusion (humanity dissolved into mush by the forces of technocapital) but not brave enough to stand against it. Yudkowsky and Alexander also started off absurdly econ-brained. But Alexander was at least willing to say that Moloch was bad, and has been gradually fishing himself out (c.f. his stuff on Schelling fences, lifeboat games, etc). Unfortunately he is less brave than Yarvin and therefore still ends up less correct on most political issues. Meanwhile Eliezer’s original framing of the alignment problem is very strongly in the paradigm of single-agent rationality. It was hard for him to then fish himself out because it feels so obvious that a superintelligence *should* be in that paradigm. But he believed deeply enough in scientific progress to lay seeds which others (especially Garrabrant and Demski) have been growing into an understanding of multi-agent rationality, which is the thing we need to reason about how to improve society. For example, he identified that economics is handicapped by assuming CDT (though he didn’t integrate that insight into his understanding of economics as deeply as Michael Vassar did).
Show more
Models will use “simulation” to justify anything, IMO it’s often motivated reasoning: (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder. In before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed. [attached is small table we ended up cutting for time but points to monitorability distinction]
Show more
🕐Announcement: is now live! A year ago, I went looking for something like a Great Replacement Tracker — a site that aggregated all the demographic data for the West and could tell me how bad things really were. To my surprise, I couldn't find one. So six months ago, I started building it myself. Today, I'm announcing the first live version of the Great Replacement Clock. It tracks the European-descended share of the world's population and of 45 historically White countries. It consolidates historical figures and modeled projections into a single interactive timeline — sourced, auditable, and methodologically transparent. Accurate ethnic demographic data for Western countries is uniquely difficult to access. No country directly tracks the share of its population that is of European descent. Several nations prohibit or restrict its collection outright; France's ban on ethnic statistics is the most prominent but not the only case. In other countries, official statistics obscure long-term trends through inconsistent census categories and aggregation methods that flatten meaningful distinctions. They miss and misclassify people in predictable ways: bundling Middle Eastern and North African populations into "White," absorbing later-generation immigrants into "native," or recording European immigrants to other European countries as merely "foreign." The US illustrates this: up through 2020, federal standards classified Middle Eastern and North African populations as White. The result is a landscape in which some of the most consequential facts about the trajectory of Western societies are among the hardest to establish with precision. This project exists to close that gap. It draws on the best available demographic evidence to reconstruct the European-descended share of each population across time, using a consistent definition across countries and eras. The sources, assumptions, and adjustments behind every figure are open to inspection. This is a v1. If you spot a questionable figure, please flag it for review. I welcome feedback and contributions. This tool will only get better with time. Demographic composition is a matter of public interest, not a protected secret. We deserve to know what time it is.
Show more
0
375
10.3K
2.6K
Forward to community