Register and share your invite link to earn from video plays and referrals.

Nathan Witkin
@NateWitkin
Research Scientist at NYU Stern's Tech and Society Lab | Tufts '24, Wesleyan '20.
1.5K Following    1.3K Followers
It's entirely fair to believe: A) discussion of extinction risks from AI is way overblown. but also: B) It'd be great if this moment leads to greater prioritization of work on alignment, interpretability, and risk reduction in the labs, government, and non-profits.
Show more
The essay below from Aaronson is just unmoored from reality, sorry. It is just not true, if you try to approach the question in a data-oriented fashion, rather than via vibes and speculative fiction and anecdotes, that capability is "dramatically ramping up ... every month." See attached graphs of progress over the previous ~year on several difficult knowledge work benchmarks. Also keep in mind that: 1) these are pass@1 scores, when what you should really be looking at to extrapolate to real world performance (esp. w/r/t full automation, "drop-in workers," etc.) is pass^n, i.e. how often, on average, models succeed at tasks n times in a row, rather than just once; 2) even strong benchmarks have low external validity because they do not and, in many cases, cannot model key constitutive properties of real world knowledge work. These include: field-specific regulatory and security constraints, ambiguous or conflicting requests from multiple stakeholders, uncodified context (related, for instance, to subtle expectations around how outputs are formatted, phrased, or otherwise presented), required use of rare or legacy software or other systems (about which there is little public data), idiosyncratic constraints on cost and timing, and sudden unexpected changes in any of the above. As I wish went without saying, any sub-tasks involved in causing catastrophic harm to humans would face many more and more difficult constraints than these. Even you if put all this to the side and just eyeball graphs, you will not see dramatic monthly ramp-ups in capability. You will certainly not see performance befitting the "machine God" that Aaronson claims is already arriving (??). As is often the case, the key error here is not to have appreciated extreme jaggedness. Solving Navier-Stokes is crazy impressive. It is also happening at the same time as models struggle to even approach human performance on many (not all!) of the complex tasks that white-collar professionals in finance, law, academia, and so on perform everyday (yes, I know NS was solved by an unreleased model; I promise that model will also underperform humans on these tasks). I can't stop you from turning your brain off and going "well obviously every form of knowledge work will fall soon, if Navier-Stokes did." But the balance of evidence suggests that this is a bad inference (even just within the field of math, by the way, where models still struggle, and will continue to struggle on many open problems). I'll begin to take doomers more seriously when they adopt a norm of trying to articulate why they expect models to best humans in fields much less hospitable to RLVR, where data is much scarcer, where tacit and context-specific knowledge plays a much larger role, and where iterative contact with the real world is essential. I'll add that it doesn't help that their concerns seem universally to pass through ill-formed concepts like RSI / AGI / ASI that they are myopically pattern-matching to a reality far too messy to accommodate them. This seemed like a point a lot of folks on here appreciated and even agreed with (esp. w/r/t AGI and ASI) not two months ago, and yet it seems to have gone out the window post-Coxon. Now these concepts have been re-drafted to serve as ineliminable premises in arguments to the effect that while models are not ready to kill us all yet they will soon pass the threshold of [key three letter premise], and so we should all be very afraid. I am not afraid and do not think you should be either. That is not because I deny AI is progressing at a rapid pace. It is instead because I see many signs that it will be both slower and more jagged than many expect, and because I see few if any signs of that progress outdoing the capacity of our institutions to adapt to it. I also find it frankly bizarre to begin worrying about extinction-level harm from AI when it can't yet hold a candle to the human cost of cars or drugs or guns or almost any other generically important technology. I see no evidence that the threat from AI is going to leapfrog past these and suddenly kill millions or billions. As a result, to organize and communicate around extinction-level threats at this stage strikes me as obviously counterproductive. All it will do and is already doing is delegitimizing the cause of AI safety (a cause I support), and creating all manner of rhetorical and strategic openings for its opponents. That is what happens when you throw millions of dollars at saying silly things you can't substantiate about a hot-button issue in public. If AI safety folks want to be taken seriously, it would help for them to be serious.
Show more
Scott Aaronson writing tonight about rats, Eliezer Yudkowsky, the Road to Damascus, and the Singularity - which he now says has already started.
No effect of AI on earnings or wages in Danish administrative data
This is revealing: "nearly every one of the Gen Z-ers we spoke to explicitly asked for systemic, top-down change to social media platforms. Not through self-control apps or screen-time reminders but through the only mechanism they believe could work: regulation." One participant drew a direct parallel to tobacco: “Before the government regulated the tobacco industry, everyone was smoking, even though they knew it was bad. Because the environment encouraged and allowed for that to happen. Yeah, the doctors would smoke. Now we need to get into a state where the environment shifts so young people’s habits shift.”
Show more
The discourse about AI is cursed because whatever you say on the topic, if you don't have to deal with exalted people who explain to you that if you don't uncritically accept their fantasies about ASI that's only because you have insufficient faith in the Machine God, you will have to deal with people who seem to believe that human intelligence is generated by invisible fairies floating in the air and that no machine will ever be able to produce anything of the sort.
Show more
This paragraph is a great example of why I find doomer reasoning unconvincing. "Solving incredibly hard problems" like Navier-Stokes and "managing massive engineering projects" are radically different! If you lump them together you are skipping over like 100 intermediate steps.
Show more
Well written essay, and I agree with much of it. But it still makes a classic AI mistake: it thinks intelligence is all you need. > These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. I think this is a category error about the world.
Show more
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc:
Show more
Now that X is interested in forecasting the growth potential of AI, it is useful to keep history in mind. According to Robert Gordon, from 1920 to 1970, electricity, the internal combustion engine, sanitation, telecommunications, and other technologies spread throughout the U.S. economy and completely changed how people live. That led to a real GDP per capita growth by about 2.46% per year. So if you expect AI to generate per-capita growth above 5%, you are effectively arguing for a technological transformation roughly twice as large as electricity, cars, telephones, modern sanitation, and larger cities COMBINED. COMBINED! Something to consider while sitting in an electrified house, after driving home in your car, browsing X on your phone from the bathroom. If you think that AI is useful, think about living without any of these.
Show more
I’ve built DNA synthesizers and sequencers by hand. I used the engineer viruses for a living. I used to engineer human immune evasion for therapeutic constructs. Who the fuck are you people? Have you ever so much as held a pipette before? You can spaghetti blast an ensemble of DNA sequences at some shitty provider but 1) you still have to assemble it and bootstrap a system for making virions 2) you are not single shotting a viable, virulent viral design without a ton of experimental selection and development. Viral fitness is deeply dependent on codons and cotranslational kinetics - you’re not just gonna obfuscate away from wild type and get something good by magic. You keep treating AI like some kinda god, but molecular physics has computational complexity that scales exponentially in particle number which just crushes the abilities of any classical computer to do end-to-end design of biological functions ab initio. Grabbing a bunch of bacteriophage phi174 hits from a mass ensemble screen in lab microbes is not evidence of some magical AGI bio design ability - it's just a classic spray and pray selection. This is just nothing like building something viable in humans. Goddamn it read some books before you waltz into biomedicine and lecture us on protecting human life.
Show more
0
76
2.3K
257
Forward to community
Many are saying
Strongly agree with Houda. Let’s retire the p(doom) thing, it’s dumb
Strongly agree with Houda. Let’s retire the p(doom) thing, it’s dumb
Every discussion about AI risk on this cursed site, summarized in a single photo.
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
Show more
0
573
14.4K
1.9K
Forward to community
I suggest simply enforcing penalties for gross negligence due to damages caused by your half-baked products. That would generate a pause without any new regulations or cartel-like agreements.
This whole narrative arc seems entirely memoryholed Perhaps more interestingly, the many indications that Anthropic staff themselves were genuinely convinced that they’d reached extremely steep RSI have also been memoryholed
Show more
This -- human understanding lagging proofs -- is the norm in biology. You get an experiment where perturbation X caused response Y, you replicate it and know it's valid, then you have to understand how it happened. Why I like the analogy of math becoming a "lab science".
Show more
Today in "tacit knowledge is a complement to AI". Caveat: the *gap* in performance using AI vs not using is higher for less skilled workers (their baseline is bad)! But the erosion of skill when trying to do the thing without AI after using AI for months is also worse for them.
Show more
Because it's infeasible to RL for. There are no sufficiently precise, context-general definitions of "funny." Fun fact, there also aren't for tons of other important things, such as "good writing," as well as many, many domain-specific varieties of what we call "taste," "judgment," etc. Failing to understand this is behind a lot of excessively bullish timelines out there.
Show more
Because it's infeasible to RL for. There are no sufficiently precise, context-general definitions of "funny." Fun fact, there also aren't for tons of other important things, such as "good writing," as well as many, many domain-specific varieties of what we call "taste," "judgment," etc. Failing to understand this is behind a lot of excessively bullish timelines out there.
Show more
I’m a broken record on this. But I think you should pay literally almost no attention to AI benchmarks that AIs and labs know in advance. To be clear, I'm very impressed by AI capabilities (to the point of being worried). But that's based on everyday use, not benchmarks.
Show more
0
37
1.1K
37
Forward to community