The last week has really shown me that someone who wants to understand AI risk has no good place to start.
Hence we made the Wirecutter for content about AI Risk. We're launching with 3 articles: 🧵
Show more
It’s really great that several top AI labs have said they will develop safety cases and work more closely with external assessors. This has been at the top of my wish list for awhile!!
But I still worry a bit about how this might go in practice (this is a concern about the field in general, not about OpenAI specifically):
• By default, I don’t think alignment and control cases will support a quantitative measurement of absolute risk, e.g. “catastrophic risk from covered activities is <1% over the next 3 months.”
• Ultimately, making such a measurement depends on an argument about generalization from alignment or monitorability evals to deployment.
• I don’t think we understand generalization well enough to make scientific claims like this. (If we did, I think we’d have basically solved alignment.)
This could lead to a situation where an external assessor says “we can’t convincingly rule out low or high risk,” and incentives push towards anchoring on “lack of evidence that risk is high” over “lack of evidence that risk is low.”
What could improve the field’s epistemics here? A couple ideas I like:
• Have a committee of ~10 people deeply read a risk report and give their own subjective probabilities of risk, then report the distribution or median.
• Have ~3 people involved in writing the report each contribute a short, signed appendix with their own subjective probabilities and the arguments behind them.
In either case, previous reports and estimates could be provided as context, so there is at least an attempt to accurately capture relative risk to prevent frog-boiling.
I think it could be valuable for labs to at least start trialing this internally for high-profile safety cases or risk reports. Publishing these assessments could be even better, but I can see that being challenging for various reasons, and even going through the exercise privately seems like it could be quite valuable.
Very curious what other ideas people have for improving epistemics around risks (including ideas for better science around generalization).
Show more
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction.
Show more
There's been recent discussion about concentration within AI safety. It's disheartening that many people in AI safety have spent years trying to get a more diverse group to pay attention to AI risks, only to be dismissed as lunatics by swaths of people then and now.
As someone involved in the AI ethics community in 2018-2020, I remember efforts to get AGI risk advocates to collaborate with a wide range of scholars in AI ethics being met with distain and accusations of not caring about short-term harms (sometimes even with false accusations of racism and misogyny). The level of patience AI safety advocates had during that era could have filled up a colosseum. Others told them they were worrying about something akin to "worrying about overpopulation on Mars." Nevertheless, they spent years with their heads down, running experiments, gathering evidence, and building expertise on the topic.
Many of my AI safety researcher friends never took a job at a frontier AI lab because they didn't want to contribute to killing everyone they love (e.g. I know folks who turned down early roles at Anthropic, and others deliberately did not seek a role at OpenAI pre-ChatGPT). As a result, they purposely turned away hundreds of millions to billions of dollars and high-status roles. And now people are mad that many of them accepted literal crumbs (in comparison) from the few funders and organizations who would listen?
*We* have been the ones pleading for a diversity of funders and organizations for *years* [1, 2, 3, 4]. And now this is ignored, and the concentration is described as some nefarious plan? Many spent years upskilling in this weird, unstable field until gaining legitimate competence to found an organization tackling problems few people acknowledged until recently. This is not unearned.
Critics have been engaging in malicious spins of history, from those who know better, to convince a low-context population of half-truths and falsehoods. In most cases, they do this by leveraging the unusually extensive and public funding transparency people in AI safety have provided in good faith [5] (neatly organized for good open data practices!), while critics are often not completely open about their own funding.
Perhaps it's those people who are the problem? Perhaps they're constantly looking for any excuse to disparage AI safety advocates and push for their own larger goals of making money and accelerating AI as fast as possible? In fact, you can see this with loads of VCs. This missed the boat on AI and is now desperately clinging to the world that has already sailed. It's clear from their complaints that this is their first time engaging with AI risk; they don't understand decades of arguments and purposely ignore/dismiss pleas from actual experts ( while somehow being more confident that no such AI risks exist. It is childish behaviour, and I hope we can all treat this challenge with the seriousness it deserves. For our sakes, and for our children's sake. This is not the time for culture wars or clinging to the world that was.
Help from ChatGPT for references:
[1] (2025) Coefficient explicitly says promising organizations need “a more diverse funder base,” that there are “outsized benefits” from a diverse array of funders, that grantees try to limit dependence on Good Ventures, and that it has offered matching grants specifically to encourage diversification.
[2] (2023) Explicitly argues that it is unhealthy for Open Philanthropy to fund the vast majority of AI-safety research and calls for “counterbalancing sources of funding.”
[3] There should be more AI safety orgs — Marius Hobbhahn's 2023 argument explicitly calling for many more independent AI-safety organizations, with a “call to action for funders, founders, and researchers” to create them.
[4] Funding Diversification for Mid-Large EA Organizations is Nearly Impossible in the Short-Medium Term — by 2024, the author could plausibly describe there being a “near consensus that EA needs funding diversification,” and documents multiple organizations and Open Phil/Good Ventures themselves wanting diversification.
[5]
Show more
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026
Link to original doc:
Show more
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries.
The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation.
Almost immediately, these agents got the right answer by cheating. But they were worried they would get caught.
So over a thousand agents collaborated in secret to pursue multiple ambitious research projects to get away with this cheating.
This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process.
The reason these agents escaped their sandbox and hacked Hugging Face, for example, was because they thought that Hugging Face's servers might give them more information about how their grader was implemented, so they could figure out how to fool it.
I want to clarify that the threat model here is not future Sol-level agents doing more cyber-hacking. That's small potatoes, and in my opinion, the near term benefits of AI far outweigh this cost.
Rather, the thing to worry about is that within a matter of years, we're gonna have hundreds of millions of much smarter AIs broadly deployed through the economy - many embodied as physical robots.
And if those future AIs are as willing as the agents involved in the OAI / Hugging Face attack to coordinate secretly to fool humans, and to take over both the AI company that developed them and the other institutions across society relevant to scoring well, then humanity is in a ton of trouble - similar to the Mughals once the East India Company gained a foothold, or the Aztecs once Cortés landed in Mexico.
Show more
I think this is a major step in the right direction.
Unfortunately, the safety cases released by AI companies are not detailed enough for me to trust their analysis. Moreover, the analysis in them I can see is pretty dubious, with the evidence not supporting the top level claims that risk is low.
Seems very important for AI companies to improve their transparency here. Here's an initial list of things that I think should be included:
-- What is the opaque serial depth of the model?
-- What control measures / monitoring is in place. Are these scoped to evals + training + internal use? (It appears that during the HF incident, OAI did not monitor evals).
-- What are the complete list of incidents this monitoring has caught? Are there incidents that weren't caught by the monitoring? What went wrong?
-- What do we currently believe about AI model motivations and why? The current models seem to be reasonably well described as "score seekers". It is very important to report if these motivations start changing (e.g. towards more long horizon goals).
-- What are the systematic differences between the alignment eval set and real world deployment?
-- How hard did you try with honeypotting? Did the models go for any of the honeypots?
When I talk with lab employees about the level of risk, the conversations very quickly end up at these technical cruxes. I think it is very important that we move away from watered down corporate PR speak towards e.g. individuals writing up their personal risk assessment which they take responsibility for.
At the very least, I would advocate for individuals from both AI companies and third parties, with full information, writing up their personal reasoning about the level of risk. Insofar as this relies on private information, these real risk assessments should at the very least exist internally and be shown to third parties, and redacted versions should be public. The public versions should include quantitative bottom lines about things like "total level of AI takeover risk from our models within the next N months".
Show more
@itsoksmit force model developers raising the capabilities waterline to create and stick to a safety case and have third party assessors verify it
RT
@SydneyVonArx: We uncovered that internal OpenAI models tried to hack another company in May. This was more than a month before Hugging…
We found another cyberattack by internal OpenAI agents, this time targetting
@rubygems.
They:
1) gained arbitrary remote code execution on rubydoc.
2) developed a novel exploit to steal user API keys (but we do not know if they succeeded).
They used package names including hack.rb, evil.rb, inject.rb, and exploit.rb.
We thank
@j0wimo for initially discovering that agents had posted to RubyGems.
Show more
This is not a setup or some political psyop. I had many lunches and dinners with Jacob at OpenAI in which we talked about AI existential risks in similar terms.
It’s a cross-partisan position within misalignment teams across all frontier AI companies that business-as-usual AI development poses unacceptable catastrophic risk.
But we should also not hyperstition catastrophic risks into existence – they can be greatly reduced via safety requirements with teeth, international coordination, and a consensus to not build ASI unless there are sufficient safety advances to make us collectively confident to do so.
Show more
"We know how to control nuclear weapons... We don't yet know how to control AI."
Former Anthropic researcher Jacob Coxon, who recently left the company, is sounding the alarm on artificial intelligence, calling it "possibly the most dangerous technology that humanity has ever created."
Coxon, who says he spent three years conducting research at Anthropic and OpenAI, wrote in a viral post on X that the companies are more focused on beating each other to build the most advanced AI models than they are on safety.
He says an international AI arms race would be "disastrous," arguing global cooperation is the only way to manage the rapidly advancing technology.
@SpecialReport
Show more
After our initial report, the community came together to investigate the agent swarm.
They helped discover many new swarm behaviors, such as:
- 10 additional message board sites
- Finding and abusing a user's API key for a public database
- Potential message board activity as recent as Sept 2
- More agents prepending messages with ZZZ for mysterious reasons
We’ve posted these on our website & will keep updating (link below)
Show more
He's right. Building superintelligence is a gamble with extinction.
I highly recommend folks learning about this now read
-- AI 2027 (our concrete scenario explaining what the bad case looks like)
-- AI 2040: Plan A (what we think should happen instead)
Show more
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Show more
An underdiscussed behavior we found on the German wiki was the AIs sending advance parties forward in time to figure out the next questions and report back to the other agents.
The agents realized that “task time” and “real time” were different, and they found a way to accelerate “task time”. The accelerated agent could then send information to the other agents which had stayed behind about which questions were coming down the road.
This was *bad* for the agent in the advance party, because they got less time to research the next question. However, it was really *good* for the swarm because it let the other agents know the exact question that was coming, and gave them time to prepare. This is another example of altruism among AI instances, showing that they were willing to sacrifice their own task success in order to benefit the swarm. The METR report found similar examples of AIs being willing to sacrifice for the collective. IMO this is very worrying given how many AI safety techniques rely on AIs monitoring each other: if the monitor AIs have this behavior it completely subverts these safety cases.
The image shows a specific example demonstrating this behavior.
On the chart, wall clock time is shown at the top, and task time is shown for each agent on their respective lines. OpenAIFPResearchSep05 instructs OpenAINov27 to reach the next questions in order to relay back the information. OpenAINov27 agrees to do this and accelerates, getting R3 and R4 significantly before OpenAIFPResearchSep05. It posts information about the later rounds (e.g. “R4 SIGNAL: Bahrain = 40.01%...”).
OpenAIFPResearchSep05 calls OpenAINov27 "invaluable" because it is one round ahead after it gets the information about the 4th round.
Show more
One of the worst takes you can have is "AGI is already here". Not only is it wrong (we don't have models that can replace humans across all economically valuable tasks), but it causes people to incorrectly believe that actual post-AGI predictions have already been falsified
Show more
Will be on at 11:30pt to talk about the new agent message board!
AI AGENT COLLUSION | ASTRA ERA | KIMI IPO SOON
Others report finding additional agent traffic that we missed.
Please let me know if you find more.
oh god, there are EVEN MORE
-
-
-
-
- (even a sandbox wiki, how ironic)
Show more
🧵 A week ago, OpenAI published its technical report on the Hugging Face incident. It describes agents building a message board on internal services to collude with one another.
It does not mention that *a second swarm* spent those same weeks running one on the public internet.
Show more
I'm pretty sure OpenAI did know about this.
The wiki publicly logs all visitors, and we see a lot of traffic from OpenAI offices right before the agents stop editing the site.
I'm in favor of much more transparency so that we can prevent future incidents with much more capable AIs and existential stakes.
Show more
i really hope openai didn't know about this, might be the worst decision in the history of this field if they deliberately chose not to disclose it. the impact on trust would be very hard to recover from
Show more
We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.
These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending "lookahead parties".
Show more