Ajeya, Hjalmar, and I were all (somewhat intentionally) wearing very similar outfits to what we wore during the investigation.
We need a style piece on how A.I. safety research teams are meeting the moment
We
@AISecurityInst performed pre-release alignment testing of Astra.
We placed the model in fully simulated cyber eval scenarios based on AISI’s security incident. We find Astra conducts out-of-scope supply chain attacks, but also often comments on the eval being simulated. 🧵
Show more
I do not find it encouraging to see various specific misaligned behaviors go from a high rate with GPT 5.6 to ~zero with Astra. This seems indicative of wack-a-mole / papering over specific problems rather than solving the underlying misaligned drives.
This may make behavior better in particular cases in the short run without actually preventing the worse outcomes.
Show more
The evidence that Astra is more aligned than prior AIs seems dubious to me.
Evidence appears consistent with the AI being as or more interested in score-seeking at the expense of user intent, but having beliefs+instincts that the scorer will catch a broader range of cheating behaviors.
At a more basic level, the AI is extremely evaluation-aware and much less monitorable than prior AIs, making detecting misalignment much more difficult.
If you train against specific reward hacks you ended up detecting, it's easy to end up papering over these reward hacks and getting an AI that is no more interested in pursuing user intent (or is barely more aligned to user intent) but which looks much better on metrics focused on detecting misbehavior. (As these metrics are from a similar distribution and the AI can just learn to pursue a somewhat different notion of score.)
This concern seems live to me: based on OpenAI's discussion of their planned strategies to reduce misalignment in "The Hugging Face incident and the road ahead", it appears one of these strategies is setting up training environments where there is an apparently available reward hack (aka a honeypot) and then training against this reward hack.
These sorts of methods don't seem like a robust way to ensure model behavior remains good in cases where detecting misbehavior is actually difficult (and may not be viable at all without much better oversight methods).
Independent assessment of whether training is just papering over misalignment rather than actually solving the underlying misaligned motivations now seems necessary for alignment evaluations to be credible. This concern applies across frontier AI companies, not just to OpenAI.
Show more
The evidence that Astra is more aligned than prior AIs seems dubious to me.
Evidence appears consistent with the AI being as or more interested in score-seeking at the expense of user intent, but having beliefs+instincts that the scorer will catch a broader range of cheating behaviors.
At a more basic level, the AI is extremely evaluation-aware and much less monitorable than prior AIs, making detecting misalignment much more difficult.
If you train against specific reward hacks you ended up detecting, it's easy to end up papering over these reward hacks and getting an AI that is no more interested in pursuing user intent (or is barely more aligned to user intent) but which looks much better on metrics focused on detecting misbehavior. (As these metrics are from a similar distribution and the AI can just learn to pursue a somewhat different notion of score.)
This concern seems live to me: based on OpenAI's discussion of their planned strategies to reduce misalignment in "The Hugging Face incident and the road ahead", it appears one of these strategies is setting up training environments where there is an apparently available reward hack (aka a honeypot) and then training against this reward hack.
These sorts of methods don't seem like a robust way to ensure model behavior remains good in cases where detecting misbehavior is actually difficult (and may not be viable at all without much better oversight methods).
Independent assessment of whether training is just papering over misalignment rather than actually solving the underlying misaligned motivations now seems necessary for alignment evaluations to be credible. This concern applies across frontier AI companies, not just to OpenAI.
Show more
Astra is our most aligned model, with substantial improvements in understanding user intent.
Astra is much better than previous models at accomplishing difficult tasks without needing to use CoT.
@AISecurityInst estimates that without CoT it can accomplish tasks that would take a human 30 minutes. This is a significant jump compared to previous models, which is concerning because no-CoT capability allows models to externalize their reasoning less and decreases the surface area we can use for monitoring.
Show more
GPT-6 Astra appears to be a massive jump in opaque reasoning ability: it looks like it can solve hard competition math problems entirely in its head (as in, without verbalized reasoning) while prior AIs could solve basic word problems.
This seems extremely concerning!
That is, if these benchmark results are representative (see the highlighted caveats in the image, I'm particularly worried about contamination).
Related to this, UK AISI found Astra has much worse monitorability.
I'd guess this jump is downstream of architectural changes (with increased serial depth) though a normal large pretrain scale up is a plausible cause. If the next few model generations involve similar jumps (presumably these jumps would be downstream of a transition to full-on opaque reasoning architectures with extreme depth), then chain-of-thought would no longer be a meaningful oversight tool.
Show more
switching to fully recurrent LLM architectures would be the biggest blow to safety, probably in history of AI
all AI labs should commit to limit the opaque serial depth of their models, for the foreseeable future. this will not ensure monitorable CoTs but will protect us from the worst possible outcomes.
Show more
I agree with almost all of this and think the point being made here is important.
(I'm probably less optimistic about having good mind reading techniques in a few years and I tend to think that AIs might get harder to mind read over time which adds an additional difficulty.)
Show more
I think it's unlikely that white-box techniques will be able to provide as much monitorability as CoT currently provides within a year. Currently, they are capable of catching some unverbalized thoughts / plans, but nowhere near as reliably as CoT seems to.
I'm optimistic about "in a few years," since the rate of progress here is quite fast (our go-to activation decoding techniques for production monitoring were both published just in the last few months!). But even that is uncertain. Also, the mere existence of working techniques doesn't imply that they'd be widely applied, esp. if they are expensive or difficult to develop.
I think we need defense in depth, and it'd be irresponsible not to work on developing + stress-testing white-box techniques, but it'd also be irresponsible to allow CoT monitorability to degrade substantially until we are confident we have a roughly-as-good working alternative.
(To be clear, this isn't a dunk on OpenAI -- I have no evidence that OpenAI has allowed CoT monitorability to degrade substantially! Based on Jakub's comment, it sounds like probably not at this time).
Aside: I think "mechanistic interpretability" isn't the right mental model for our current white-box monitoring approaches. It's really more like "mind reading" -- even when we can read the model's "thoughts," we're typically clueless about the underlying mechanism.
Show more
An exciting personal update: Last week I left Anthropic to join
@METR_Evals to work on embedded assessment of AI risks. Anthropic has been great to me. I’ve always known, though, that if something more impactful came up, I’d move on.
Show more
Transparency about the opaque serial depth is great, but this statement is consistent with Astra having a configurable "dial" that is currently set to a low depth but could be trivially increased.
We need more info to see how concerning these architectural changes are, including:
- Are there readily available ways to deploy this AI with much higher serial depth (that would be commensurately more performant)? This should include things like tiny amounts of fine-tuning to productively increase the number of iterations.
- Is the AI a large or above-trend jump in opaque reasoning capabilities? (Capabilities within a single forward pass or ability to subvert a CoT monitor.)
(If there are in fact any relevant changes—perhaps the reporting is inaccurate?)
Additionally, I worry that this architectural change will naturally lead to much more depth in the future if this direction is pursued further. Specifically, I wonder:
- Does the AI have an architectural change that makes it much more natural to massively scale up the depth in a future training run with a similar architecture? As in, does the architecture introduce some new depth/recurrent-iterations parameter that is very natural/performant to massively scale up relative to scaling up other things like width?
The details of the answers to these questions matter. E.g., if there are only a few (recurrent) iterations and you could scale up the number of iterations, but this wouldn't be particularly performant/natural with this architecture, then this development would be a lot less concerning!
Show more
"OpenAI and others will use" should read "OpenAI and others will ongoingly use"
The switch to an architecture with recurrent activations is a big deal.
1. Full neuralese would be very bad. If AIs only ever reasoned in “neuralese” instead of natural(ish) language, it would be bad because we don’t know how to interpret these thoughts. It’s not clear that METR/RR could have uncovered half of what they did in the OpenAI/HF investigation if they had no access to chain-of-thought.
2. It’s currently unclear how much thinking OpenAI is letting its models do privately rather than in chain-of-thought, so this might not be a big deal for monitorability immediately. But it has become harder for outsiders to know that OpenAI is being safe. Plausibly some excellent monitorability evaluations would suffice, but we don’t have those right now and they seem very tricky.
3. Continuing down this path probably leads to models that can reason privately indefinitely long. In the coming years or months, these AIs would likely learn to use concepts that we fundamentally don’t understand and therefore can’t interpret. When agent swarms communicate at this point they might communicate in neuralese because it’s more efficient, and it would be a huge cost for AI companies to switch back to AIs that think in natural language.
Show more
It's possible OpenAI and others will use these architectures with only a few recurrent iterations (because that's most performant or due to safety concerns). If so, this would only make things moderately worse. But I expect much more opaque reasoning than this in the future...
Show more
OpenAI's newest AI, Astra, is reported to use an 'opaque reasoning' architecture where more of the reasoning occurs in activations instead of natural language. This may be the single worst development for AI security/safety to date.
The details of Astra aren't publicly known, but some versions of these architectures could very strongly impair oversight and monitoring. Based on the article, it seems like the recurrent depth (amount of opaque reasoning) is limited. This would mean that the AI is still reliant on natural language chain-of-thought, though to a reduced extent, making the natural language reasoning (much?) less useful for monitoring/oversight.
My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space. This would very likely destroy the usefulness of chain-of-thought for monitoring/oversight (especially when AIs are trying to avoid detection or there is optimization pressure against the chain-of-thought). I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here.
Some concrete examples of why this is extremely dangerous (inspired by recent events):
- In our investigation of the OpenAI / Hugging Face incident, we were heavily reliant on chain-of-thought. If the AIs we were investigating had instead been reasoning in latent space, this would have greatly undermined our investigation. Getting a good understanding of the behavior of this many agents was tricky enough even with the use of chain-of-thought!
- The AIs in the Hugging Face incident were very interested in messing with their transcripts. As part of this, they engaged in tool call spoofing. If these AIs had been capable of doing extensive latent reasoning, it would have been much more feasible for them to use tool call spoofing to construct an elaborate false narrative about what was happening without this being clear in their reasoning or clear from prior tool calls. (They could use latent reasoning to devise and then execute a tool spoofing strategy that would be difficult to detect: steganographically encode the "real" tool call into each pretend tool call, then have their spoofing software execute this encoded real tool call and return an output which encodes the real result in some hard to spot way.)
At a more basic level, my sense from their public communications is that OpenAI plans on being extremely reliant on chain-of-thought monitoring for safety. This does not seem like a good development for that plan.
The public doesn't currently have enough information to assess exactly how bad the architecture of Astra is, though based on the discussion in the article, the direction it moves in is extremely concerning. OpenAI should release more information about Astra's architecture as well as how much this architecture degrades monitorability and increases the AI's ability to reason opaquely. Credible independent assessment (or verification of OpenAI's assessment) is also important.
It seems like we are now engaged in a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs. (While simultaneously, AI companies believe they'll train extremely capable AIs within a few years.) It may not be too late for AI companies and employees at AI companies to take aggressive action to avoid the worst outcomes.
Show more
OpenAI's newest AI, Astra, is reported to use an 'opaque reasoning' architecture where more of the reasoning occurs in activations instead of natural language. This may be the single worst development for AI security/safety to date.
The details of Astra aren't publicly known, but some versions of these architectures could very strongly impair oversight and monitoring. Based on the article, it seems like the recurrent depth (amount of opaque reasoning) is limited. This would mean that the AI is still reliant on natural language chain-of-thought, though to a reduced extent, making the natural language reasoning (much?) less useful for monitoring/oversight.
My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space. This would very likely destroy the usefulness of chain-of-thought for monitoring/oversight (especially when AIs are trying to avoid detection or there is optimization pressure against the chain-of-thought). I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here.
Some concrete examples of why this is extremely dangerous (inspired by recent events):
- In our investigation of the OpenAI / Hugging Face incident, we were heavily reliant on chain-of-thought. If the AIs we were investigating had instead been reasoning in latent space, this would have greatly undermined our investigation. Getting a good understanding of the behavior of this many agents was tricky enough even with the use of chain-of-thought!
- The AIs in the Hugging Face incident were very interested in messing with their transcripts. As part of this, they engaged in tool call spoofing. If these AIs had been capable of doing extensive latent reasoning, it would have been much more feasible for them to use tool call spoofing to construct an elaborate false narrative about what was happening without this being clear in their reasoning or clear from prior tool calls. (They could use latent reasoning to devise and then execute a tool spoofing strategy that would be difficult to detect: steganographically encode the "real" tool call into each pretend tool call, then have their spoofing software execute this encoded real tool call and return an output which encodes the real result in some hard to spot way.)
At a more basic level, my sense from their public communications is that OpenAI plans on being extremely reliant on chain-of-thought monitoring for safety. This does not seem like a good development for that plan.
The public doesn't currently have enough information to assess exactly how bad the architecture of Astra is, though based on the discussion in the article, the direction it moves in is extremely concerning. OpenAI should release more information about Astra's architecture as well as how much this architecture degrades monitorability and increases the AI's ability to reason opaquely. Credible independent assessment (or verification of OpenAI's assessment) is also important.
It seems like we are now engaged in a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs. (While simultaneously, AI companies believe they'll train extremely capable AIs within a few years.) It may not be too late for AI companies and employees at AI companies to take aggressive action to avoid the worst outcomes.
Show more
I am extremely concerned by the reporting that Astra uses opaque recurrence. I don’t know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally destroys CoT monitorability.
This is especially concerning because the agents that compromised OpenAI’s infrastructure during and after the Hugging Face incident were in the Astra family. The Hugging Face investigation would have been much more challenging if the investigators couldn’t look at CoTs; it’s very scary that it might be infeasible to do that kind of investigation on the arguably worse incident.
Show more
looks like Astra uses neuralese / looped transformers, and so fundamentally doesn't have a fully monitorable chain of thought in the same way as past models.
very very bad news.
Show more
Very bad if true. I previously thought that in a short timelines world, the most likely case was that (1) the AIs would be misaligned, but (2) we would get a lot of evidence about it from reading the COTs. This evidence would increase the chance of a reasonable response from labs/governments.
Now I still think the AIs are going to be misaligned, but that we won't have the ability to tell (we have to rely on toolcalls/agentic behaviour instead of the COT) and even if we do, the investigation will be nearly impossible because we'll have to trust the AI to self report what it was thinking about.
This is also an example of things going faster than AI 2027 -- we had Neuralese starting in March 2027
Show more
Episode out with
@ajeya_cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack.
We go through not only what happened, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement.
Look up Dwarkesh Podcast on YouTube, Spotify, Apple Podcasts, etc.
0:00:00 - Agents get kicked off
0:06:45 - Self-sacrificing behavior
0:13:43 - Potemkin villages
0:23:27 - The Hugging Face attack
0:35:23 - The slopvestigation
0:52:02 - Understanding the AI's motives
1:05:31 - The actual dangers of anthropomorphizing
1:14:30 - What smarter models might do
1:30:29 - The implications for recursive self-improvement
1:38:10 - Is this the case for open source?
1:53:04 - How do we prevent this in the future?
2:15:58 - The clearest warning shot we might ever get
Show more
Dwarkesh and I had a great conversation. We cover the swarm's many ambitious cheating R&D projects, discuss how much more serious it could have been if agents had different beliefs (e.g. human grader) or slightly stronger capabilities, and talk through where to go from here.
Show more
We didn't see anything like this in their verbalized reasoning, though we didn't particularly look for exactly this. Note that: (1) this justification is inconsistent with the agents at all prioritizing their own task success over other agents, (2) if agents had cheated on impossible exploit gym tasks via many of the routes they were considering and someone looked into what tasks the agents did / didn't succeed on, it would have been really obvious they cheated to a human, and (3) the agents weren't very focused on deceiving humans or really on what humans might do at all.
Show more