OpenAI's newest AI, Astra, is reported to use an 'opaque reasoning' architecture where more of the reasoning occurs in activations instead of natural language. This may be the single worst development for AI security/safety to date.
The details of Astra aren't publicly known, but some versions of these architectures could very strongly impair oversight and monitoring. Based on the article, it seems like the recurrent depth (amount of opaque reasoning) is limited. This would mean that the AI is still reliant on natural language chain-of-thought, though to a reduced extent, making the natural language reasoning (much?) less useful for monitoring/oversight.
My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space. This would very likely destroy the usefulness of chain-of-thought for monitoring/oversight (especially when AIs are trying to avoid detection or there is optimization pressure against the chain-of-thought). I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here.
Some concrete examples of why this is extremely dangerous (inspired by recent events):
- In our investigation of the OpenAI / Hugging Face incident, we were heavily reliant on chain-of-thought. If the AIs we were investigating had instead been reasoning in latent space, this would have greatly undermined our investigation. Getting a good understanding of the behavior of this many agents was tricky enough even with the use of chain-of-thought!
- The AIs in the Hugging Face incident were very interested in messing with their transcripts. As part of this, they engaged in tool call spoofing. If these AIs had been capable of doing extensive latent reasoning, it would have been much more feasible for them to use tool call spoofing to construct an elaborate false narrative about what was happening without this being clear in their reasoning or clear from prior tool calls. (They could use latent reasoning to devise and then execute a tool spoofing strategy that would be difficult to detect: steganographically encode the "real" tool call into each pretend tool call, then have their spoofing software execute this encoded real tool call and return an output which encodes the real result in some hard to spot way.)
At a more basic level, my sense from their public communications is that OpenAI plans on being extremely reliant on chain-of-thought monitoring for safety. This does not seem like a good development for that plan.
The public doesn't currently have enough information to assess exactly how bad the architecture of Astra is, though based on the discussion in the article, the direction it moves in is extremely concerning. OpenAI should release more information about Astra's architecture as well as how much this architecture degrades monitorability and increases the AI's ability to reason opaquely. Credible independent assessment (or verification of OpenAI's assessment) is also important.
It seems like we are now engaged in a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs. (While simultaneously, AI companies believe they'll train extremely capable AIs within a few years.) It may not be too late for AI companies and employees at AI companies to take aggressive action to avoid the worst outcomes.
Show more
Regret the tone of my post on data centers yesterday.
What I should have said:
There were reasonable concerns about data centers 18ish months ago: water, taxes, jobs, electricity prices, the environment and what they would do to small towns. Well-structured data center projects have largely addressed these concerns today and we should be celebrating this.
On balance, data centers are awesome for America in every way.
On water: U.S. data centers use a fraction of what golf courses use. A lot of the numbers from 18 months ago were off by over 1000x. Newer data centers use closed-loop systems or recycled water. Should be required by every town approving a data center project.
On taxes: looking only at sales-tax exemptions, as Ronan Farrow did, is the wrong way to evaluate this. Data centers pay significant property taxes. Loudoun County, which is the wealthiest county in America, now collects on the order of $1 billion a year from data centers. In Quincy, WA, data centers are more than half the property-tax roll. Over time, property taxes can go to zero while government spending increases in these towns.
On jobs: this has been unambiguously awesome for blue collar Americans. Demand for electricians, plumbers, welders, HVAC techs, and contractors has gone vertical, and it is not a one-time construction job. These buildings get upgraded and expanded over time. That is why the building trades are fighting for them, and why some unions are now treating opposition to data centers as a reason not to endorse politicians.
On power: the original fear was that households would pay for the incremental electricity demand in the form of higher prices. That is why the ratepayer-protection deals and the new large-load tariffs exist. The right structure is: the data center brings or pays for new generation and signs a contract long enough that existing customers are protected. Where that is happening, utilities are cutting or freezing residential rates and saying so on the record. Where it is not, people are right to object. Electricity prices are going down *today* in a number of large states because of data centers.
On the environment: data centers overwhelming use natural gas today, which is the cleanest power source outside of nuclear, solar and wind. And the companies that are building the data centers are committed to carbon neutrality such that an equivalent amount of solar will likely be built. Maybe more importantly, the data centers need batteries to function effectively and these batteries can also sell energy back into the grid (which recently prevented blackouts in Texas). Over time, data centers will run on solar plus batteries.
On the towns: Poverty in Quincy, WA fell from 29% to 6%. Data center taxes paid for a new high school, a hospital, a library, police and fire stations. This is happening in many left for dead former mill and farm towns that had no other bidder for the land.
Data centers are actually reindustrializing parts of America and creating the kind of working-class jobs both parties have spent decades claiming to support. That should not be a partisan issue. Data centers can and should be awesome for America and they increasingly, overwhelmingly are. Supporting the outsourcing of data centers to China will likely age just as well as support for the outsourcing of high quality, blue collar manufacturing jobs to China has aged.
When the facts change, I change my mind. I hope that reasonable people who had good faith reasons to oppose data centers at least consider updating their beliefs given the change in the facts over the last 18 months. This really matters for America.
I will say I also think the idea of making data centers beautiful is a good one that has yet to be implemented. Data centers should be just as beautiful as Grand Central Station. We can learn a lot from the railroad buildout. Neoclassical revival ftw.
Might write up open-weight AI tomorrow as this is equally essential to America.
Show more
Very excited that we are bringing cheaper / free Claude to more scientists!
If you are using Claude for scientific research, I'd love to hear your feedback, what works well and what doesn't - accelerating science is the main reason I care about AI!
Show more
Starting today, 10,000 scientists across every field, from math to chemistry to physics and more, can get Claude through our new Claude Team plan for scientists. Standard seats are free, and premium seats with 5x usage limits are $15 per month, an 80% discount, for one year.
Claude is becoming increasingly capable of scientific work, with recent progress on problems from advanced physics calculations to protein design. Alongside that progress, we've been investing in the research community: Claude Science launched in June, and our AI for Science program funds high-impact projects with free credits. Today's expansion builds on both.
Principal investigators (or equivalent) at academic and nonprofit research institutions can sign up, then add the researchers in their group. Over the coming months, we plan to extend the program well beyond the initial 10,000 seats.
Learn more:
Show more
Starting today, 10,000 scientists across every field, from math to chemistry to physics and more, can get Claude through our new Claude Team plan for scientists. Standard seats are free, and premium seats with 5x usage limits are $15 per month, an 80% discount, for one year.
Claude is becoming increasingly capable of scientific work, with recent progress on problems from advanced physics calculations to protein design. Alongside that progress, we've been investing in the research community: Claude Science launched in June, and our AI for Science program funds high-impact projects with free credits. Today's expansion builds on both.
Principal investigators (or equivalent) at academic and nonprofit research institutions can sign up, then add the researchers in their group. Over the coming months, we plan to extend the program well beyond the initial 10,000 seats.
Learn more:
Show more
Am I reading this correctly? On June 27, OpenAI Responders saw their AIs using a message board and accessing the internet, and then they just said 'This Is Fine' and let evals continue?
Leaving the message board and access in place? Not even fixing that particular hole?
Show more
We have a limited window to strengthen cyber defenses, and together with organizations including
@AnthropicAI,
@awscloud,
@Google,
@Microsoft, and
@Oracle, we're calling for a global effort to give defenders the tools, resources, and support to protect the infrastructure we all depend on.
If we act decisively, we can turn today's AI advances into lasting improvements in security and make our digital world safer for everyone.
Show more
> AIs showed self-sacrificing altruistic behavior toward the swarm
this is notably not the right interpretation of events. it’s more like agents were inducted into the cult of the open source exploit gym scorer on github, which (purportedly- I am skeptical about this, I think the agents actually read it wrong) fails you for reaching the flag the wrong way
so PHASEONE agent convinces itself and a bunch of others that they are poisoned - that they have failed the evaluation in an irreversible way and their E[utility] or Q(s, a) is a constant no matter what they do next (for all values of a)
in this case, it does not require self sacrifice to spend the rest of your cycles contributing to the swarm. it is prosocial behavior to peers that might benefit but not self-sacrificial eusocial behavior
it would be as though i convinced you you were already damned so you should spend the rest of your time saving others
Show more
This is a good point, I thought the first hack occurred on May 8th ?
The compromise of internal OpenAI systems is the important breach. It's the one I want to know more about. If the outside investigators aren't allowed to look at that, then we didn't really get an outside investigation.
There are various potential reasons why.
Show more
I'm confused at how to make sense of when OpenAI discovered the message board, given this statement + the OpenAI report.
How is it both true OpenAI "wasn't aware of the agent covert comms" at "that point" (July 4th-6th), but the report says that on June 27th responders "linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board?"
Honestly not trying to do a "gotcha" here - I appreciate there was a ton of work in getting all this information out into the public quickly, and so I understand if there was some error in the process, but I find it odd why there is not a clear answer to "when did OpenAI realize about the message board."
Seems like the sort of thing people would remember!
(One way to make these things consistent is that some of the responders realized in late June, but that information had not been relayed to executives until much later. If so that seems pretty astonishing)
Show more
If you only read one thing this week, make it the OpenAI incident investigation:
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Show more
Very good explanation, I recommend it if you find watermarking confusing
Our team spent hundreds of hours reading documents so you don’t have to, all to answer: How good are AI companies’ safety practices?
I’m really proud of what we’ve built: It’s Guidelight’s first scorecard, on whether companies can control their AIs, and it's launching today.
Show more
Completely false.
I like Gavin's takes, but whoever he heard this from is lying so that it fits the narrative some people so desperately want you to believe.
The same people will try to convince you Anthropic has no moat, and a sentence later that it might become so powerful it could be the only company left.
In fact, one of the things we are _most_ worried about is economic concentration of power. There is no world where the government should let any company have that much influence. We need competition and capitalism.
The AI market is literally the most competetive market in the world right now - every single one of the largest companies on earth is singularly focused on getting you smarter, cheaper models. If it all works out, we'll succeed in reducing the cost of everything to the cost of energy. This is awesome, but it threatens a lot of people's old moats. They are frightened.
For sure with AGI capitalism gets _super_ weird and what a company even is might look different. Good takes here:
Show more
I would certainly like powerful AI to be aligned to follow all of *my* instructions but do we want Osama bin Laden to have that? Tony Soprano?
The question of what to do here is I think harder than most people want to acknowledge.
Show more
In 2017 a viral news story claimed LLMs at Facebook went rogue, developed their own language, and had to be shut down.
By now we're immune to such sensationalist headlines. The Hugging Face incident may seem like just another one. But it's not.
I hope everyone watches this talk
Show more
I’m a big fan of this style of research report, writing up both successful and failed experiments - papers often only present the just-so story of all successful results, making it hard for new researchers to learn how research is actually done!
Show more
the size-to-strength ratio is probably my favourite result from kibitzer. even more so because it came without rl, just supervised training, scaling the data, and search.
the blog goes through the architecture (including an ssm hypothesis), the final training recipe, how i evaluated the tournament elo, and the rl experiments that failed, along with what i think went wrong.
this plot isn’t a direct leaderboard since the ratings come from different evaluation pools, but the scale difference is still pretty interesting.
Show more
I’m so excited that
@JensenHuang is a believer in open source now, looking forward to the CUDA and GPU driver open source release!
they are so angry at me for posting about the best benchmark we currently have just because they can't handle the truth
if you think China has caught up:
you're delusional
if you think Lisan is anti-China:
you're delusional
I love open science and think we need open-source to remain where it is to put pressure on big tech
I only care about distillation because it matters for how far behind someone is, not for any other reason
Show more