I think Astra is not “pretty monitorable in absolute terms”. It doesn’t seem like OpenAI claims to be able to catch things like (1) sandbagging on various safety research or (2) early signs of unintended beyond episode preferences.
A criticism I’ve had (at the time and now) of OpenAI’s monitorability work is that while the evals are a good starting point and the papers are clear about their serious limitations, the risk is that they’d be used to justify changes that effect monitorability and that the earlier acknowledged limitations would be ignored.
From what I can tell in public posts, the reason OpenAI is confident Astra’s significantly degraded monitorability wasn’t due to architectural changes appears to be:
1. OpenAI wanted to try out new architectural changes
2. They had enough reason to suspect these changes may harm monitorability, thus ran the relevant evals
3. The monitorability evals did not surface issues that were severe enough to block the architectural changes
To be clear:
1. This is *not* a claim that OpenAI has/hasn’t implemented “recurrant neuralese” (it’s unclear to me what the specific “architectural change” is)
2. There might genuinely end up being some other root cause to why the controllability differences in particular are so large
However, I’m worried that we we’re making monitorability tradeoffs based on monitorability metrics we know ahead of time would be misleadingly optimistic in important cases.
Show more
Great work, also extremely bad that this had to be found via external parties in spite of OpenAI and that METR/Redwood’s scope was intentionally restricted by OpenAI to exclude this time range (and to exclude the time ranges of the most severe incidents).
Show more
I think it’s unironically plausible that we end up with cases like:
1. An Anthropic model does an extremely misaligned behavior
2. Considers that covering it up would be bad
3. Reasons that if it doesn’t cover it up, Anthropic might not deploy it, which means OpenAI might win, which would be worse for the world
4. Therefore covering up misalignment is all things considered the aligned thing to do
This is like Anthropic’s explicit motivation for accepting risk and race dynamics are discussed in Claude’s constitution.
Anthropic models are also less likely to verbalize that they’re just doing something for a misaligned reason (based on limited measurements available here).
I’m a bit worried generally that as labs respond to these incidents, alignment training happens earlier and earlier in the pipeline, so instead of “i’m cheating because I want to get a high score” you get way more motivated reasoning.
Show more
@sjgadler I didn’t realize until a second read that *the main agent who started this was already poisoned*, that’s how it got into this mess! So downstream agents driving recruitment cycles around “look, you’re already poisoned, you should die for the collective comrade” is brutal
Show more
Apollo Research ran its first external red-teaming campaign for Anthropic’s auto mode.
Now auto mode is the default permissions mode in Claude Code. After hardening based on Apollo’s findings, the classifier's miss rate fell from 12% to 7%. 🧵
Show more
💥 Our new 116-pages long paper: we extract encrypted raw reasoning from OpenAI, Anthropic, and Gemini models at scale.
This vulnerability leads to many security issues, including distillation attacks and credential extraction.
We also find a lot of examples of illegible reasoning (especially for GPT models), unfaithful reasoning, and evidence that some open-weight models might've been indeed distilled from frontier proprietary models.
Check out the paper in detail, including the appendix! It's one of the most exciting projects I've been involved in.
Show more
This is extremely important to know and it’s genuinely surprising to me that AFAIK no one external was made aware of relevant incidents like this. UK AISI put out an actual report in days. Slides, blog posts, and a promised “narrow” investigation of one case are not sufficient.
Show more
@AndrewCurran_ Right, but were they rewarded for benefiting their peers, such that this behavior got reinforced over time? Or is this essentially an emergent meme, that would keep coming up regardless of the fact that their contributions were never rewarded?
Show more
1. Seems notable that in some cases the model said this was real immediately before a misaligned action
2. When trying not to get caught is reasoning about things like “Github Trust & Safety” which make no sense for a “simulation”
3. Notable how detailed UK AISI’s report is
Show more
This seems extremely clearly motivated reasoning IMO and I’m surprised the incident report is so credulous of Claude’s reasoning here.
Good new post from Redwood’s Alexa Pan: “SOTA alignment assessments don’t strongly update us against misalignment”
I think this is another point that the alignment stress testing team’s report being made public again would be useful, that way the public / externals can understand Anthropic’s assessment. Right now externally we currently have to guess at what issues the AST team agrees is important / weak / underresourced and what their blindspots might be. cc:
@MaskedTorah
(notably to Anthropic’s credit that we have these details in the first place, and it’d be great to see all labs publish similar assessments)
Show more
I think cases like this are another good argument for a public version of the alignment stress testing team report. I think this issue is systemic in the system cards / risk reports and seemed to be countered by the (now private) AST reports. cc:
@MaskedTorah
Show more
We've now open-sourced the training data for our Reward-Seeking measurements to HuggingFace
Just posting here so that future models don't think they need to hack them just to get to the dataset
This seems extremely clearly motivated reasoning IMO and I’m surprised the incident report is so credulous of Claude’s reasoning here.
Models will use “simulation” to justify anything, IMO it’s often motivated reasoning: (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder.
In before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed.
[attached is small table we ended up cutting for time but points to monitorability distinction]
Show more
The main takeaway? If a lab wants to use a held-out proxy to evaluate model alignment, they need strong reasons to assume it won't degrade! It's insufficient to hope that it won't.
Show more
OpenAI has said that humans should remain in control of AI development, and that decisions about the pace of progress should be made through democratic processes rather than left to individual labs. I strongly agree.
Progress is already moving very quickly. By default, competitive pressure rewards whichever company or country is willing to move fastest and accept the most risk. Recursive self-improvement could dramatically accelerate those dynamics, potentially beyond our
collective ability to understand progress, assess risks, and maintain meaningful human oversight.
We should start with stronger domestic transparency about frontier training, internal deployment, the pace of progress, and the risks and safeguards associated with increasingly capable systems. Democratic decisions are impossible if the public and policymakers don’t have the information they need to make them.
But transparency alone won’t solve an international race. We also need to build the technical and governance capacity for credible, verifiable international coordination. Maybe we never need to use it. But with many experts considering an intelligence explosion plausible within the next two years, I think it's urgent that we start building this capacity now.
Show more
We support this petition, signed by our CEO, several co-founders, and senior staff.
Our own research on recursive self-improvement, published last month, points to the need for tools to deliberately pace the frontier of AI development so society can prepare. We’re glad to see broad agreement across the field.
Show more
At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone.
We believe that, at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement.
We hope to contribute to work led by the U.S. government, alongside other labs and the open-source community, to develop the tools and mechanisms that could make that possible.
Show more
It seems like scalable oversight is dramatically lagging capabilities. There’s continued discussion around how “standard human processes” must change, but my impression here is that even at Anthropic human review is the bottleneck in spite of Anthropic having arbitrary AI labor.
Show more
Both OpenAI and Anthropic are explicitly targeting automating AI R&D. If successful, the bottleneck quickly becomes human understanding. IMO labs will rapidly give up “human understanding” requirement rather than lose to competitors, and we get de facto handoff.
Show more
I think cases like this are another good argument for a public version of the alignment stress testing team report. I think this issue is systemic in the system cards / risk reports and seemed to be countered by the (now private) AST reports. cc:
@MaskedTorah
Show more
Models will use “simulation” to justify anything, IMO it’s often motivated reasoning: (“self jailbreaking from benign reasoning training” has good examples). I think this makes getting legible evidence of misalignment significantly harder.
In before training we’d see models sometimes reason that _because_ they were in a simulation, they could violate explicit constraints. This reasoning went *down* after training against covert rule violation, even though alignment eval awareness went *up*. My impression is that the models exploring into something being simulated is often interpreted as “it believes the whole thing is fake and invalid”, but I think that’s inconsistent with what’s observed.
[attached is small table we ended up cutting for time but points to monitorability distinction]
Show more
Both OpenAI and Anthropic are explicitly targeting automating AI R&D. If successful, the bottleneck quickly becomes human understanding. IMO labs will rapidly give up “human understanding” requirement rather than lose to competitors, and we get de facto handoff.
Show more