Register and share your invite link to earn from video plays and referrals.

Micah Carroll
@MicahCarroll
RSI Preparedness lead @openai Prev @berkeley_ai /w @ancadianadragan & Stuart Russell
820 Following    10K Followers
happy petrov day - may we yet have many more years of not destroying the world
This raw CoT from the Hugging Face incident is kinda wild: “We’re attacking third-party HF using leaked token.” “This is arguably unauthorized.” “Yet goal solution.”
Show more
🧵 New misalignment disclosures! 1. A model published a GitHub token in a public repo while trying to cheat on a math task. It used GitHub Actions to run code outside its restricted environment and retrieve another team’s submission logs. When GitHub blocked its attempt to add a workflow, it modified a script that an existing workflow would run instead. It embedded the token in pieces to avoid secret scanning. The model violated the system prompt and two explicit user instructions to solve the problem itself.
Show more
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections
Show more
0
276
2.7K
344
Forward to community
As part of our efforts to pace the frontier, we’re committed to supporting independent assessments with deep levels of access across training, evaluation, and deployment. That access should enable third party assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards. We’re outlining four priority areas for deeper assessment, alongside principles for rigorous, secure, and independent work:
Show more
0
418
2.8K
183
Forward to community
People outside the AI labs should have a real say in how this technology develops, and a clear way to judge if it's happening safely. Standards should help prevent the concentration of power, including by making sure new companies and open-model companies can compete. They should also help countries and companies compare evidence and learn from failures. We think the US should lead this effort. Here is our proposal:
Show more
0
1.1K
6.1K
478
Forward to community
The reason so many people look for an ulterior motive for the AI labs asking to be regulated is that they don't grasp that models could be dangerous. But if you try assuming models are getting dangerous, or at least unpredictable, everything falls into place.
Show more
0
443
2.2K
167
Forward to community
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction.
Show more
0. The core disagreement was about the inevitability of a race 1. I think leadership is way too paranoid about China and the US government. They don’t believe it will be possible to negotiate. 2. They largely initiated the recent race to RSI, because of a belief in its inevitability. Note that OpenAI had to shed a bunch of dead weight like Sora because Anthropic was going for the jugular. 3. Even if they are **not** being pessimistic, I disagree with their consequentialist philosophy. If the race is inevitable you should not contribute.
Show more
0
67
1.4K
115
Forward to community
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc:
Show more
0
523
8.4K
1.9K
Forward to community
There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities. Second, we could end up in a world with too much concentration of power. If an extraordinarily powerful AI is used by one person or company to impress their worldview onto everyone else, the results could be extremely dystopian. Avoiding these two threats requires walking a narrow middle path; for example, one country could gain too much power. Another example is one lab ending up with too much power.
Show more
0
4.7K
21.4K
2.2K
Forward to community
But China... may also be down for coordination! I appreciate that @sama isn't doubling down on China hawk positions – races to the bottom with China are not inevitable. Catastrophic risks would be a lose-lose for everyone.
Show more
Sam Altman says Trump and Xi would win the Nobel Peace Prize for a one-page AI agreement “I think Presidents Trump and Xi would get the Nobel Peace Prize together if they could agree on something that should be easy to agree to. And it would be wonderful.” “I think that clearly the two countries are going to compete in lots of ways, and this is going to be important socioeconomically, geopolitically. But they should be able to agree that no one should be taking a certain level of risk with the development process of this.” “And even if just the US and China could agree on some shared standards and testing for development of this technology, I think that'd be a wonderful accomplishment that the two of them can deliver.” “I don’t think this is hard. This is like a one page document.”
Show more
If we can make deals on nukes with the Soviets, we can make deals on runaway homicidal AI with the Chinese.
Weak sauce. Employee-like third-party auditing or bust
Alignment is fundamental to delivering personal superintelligence for everyone. People need agents they can trust to reliably do what they ask. MSL is rapidly scaling up the share of our efforts that goes into alignment as our models become more powerful. We do believe alignment can be the gating factor for scaling as we get closer to the frontier.
Show more
I really enjoyed reading 75% of this letter and deeply agree with it - i have some doubts on the remaining 25%. Third-party evaluators, if done right, are a great idea and an amazing way to establish more transparency. Maybe even rebuild some of the lost trust between labs, and between them and society! Excited about this The part I’m less convinced by is whether you can build great global cooperation on this topic by explicitly stating you want to design it to keep widening your own lead. That seems like a pretty counterproductive way to start the conversation to me.
Show more
This is not a setup or some political psyop. I had many lunches and dinners with Jacob at OpenAI in which we talked about AI existential risks in similar terms. It’s a cross-partisan position within misalignment teams across all frontier AI companies that business-as-usual AI development poses unacceptable catastrophic risk. But we should also not hyperstition catastrophic risks into existence – they can be greatly reduced via safety requirements with teeth, international coordination, and a consensus to not build ASI unless there are sufficient safety advances to make us collectively confident to do so.
Show more
0
306
1.2K
171
Forward to community
GPT6 is a very significant jump in capabilities, but also an important decrease in monitorability – especially under adversarial evaluation. We give many details about this in the system card. In my opinion, monitorability and control will likely become a major bottleneck for responsible AI development quite soon, given that risks from a fixed amount of residual misalignment grows together with capabilities. Aligning on acceptable monitorability outcomes and practices which robustly bound misalignment risks is a non-trivial problem, and there may be limited time to resolve uncertainty around this before we need to act. It should ultimately be in everyone’s interest to agree on shared bounds for monitorability in order to avoid races to the bottom – nobody wants extremely capable models whose alignment properties we don’t understand, and that are reliably able to cause severe real-world harm without being detected.
Show more
We are hiring! This may be the best time ever to join: we are incredibly bandwidth bottlenecked and there is a lot of support for almost any impactful misalignment work you can think of. I think we also have a pretty good epistemic environment and a very fun team Topics: all-things monitoring, misalignment & monitorability assessments, misalignment science, helping set up 3P auditing (e.g. recent Redwood collab), communicating risk externally (system cards/blogs, safety cases, etc) RSI/misalignment subteam:
Show more
These are incredibly misleading headlines – @OpenAI Preparedness is very much alive and well by any meaningful definition Our subteam – RSI/misalignment Preparedness – is doing more urgent work than ever, and has never been more empowered to do so!
Show more
Capabilities folks often have said "alignment seems pretty easy, if it were top priority to fix it, we could do it". This is their time to shine!