I just finished watching this and it was another excellent episode that I strongly recommend if you're interested in the topic.
Many people have criticized Brown for engaging in hyperbole because of the example he used to illustrate how difficult it is to be confident that isolation is air tight, but 1) he is actually very good at conveying the uncertainty about how AI will progress from this point on and 2) the point he was making in that part of the discussion was not unreasonable in context and I don't think the way in which people interpreted it is charitable.
As someone who constantly bitches about how people are overconfident, I wish there were more discussions like this where people take uncertainty seriously.
Show more
A few thoughts on this:
1) If you’ve only seen clips of this interview, I’d encourage you to watch the full podcast. I push back on plenty of AI hype in it.
2) As I said in the podcast, this example is academic. My intention was to illustrate how hard it is to make absolute guarantees about isolation, which is why it's important to have layers of defense. The part before the clip starts is me talking about other layers of defense.
3) The example I'm bringing up isn't about weight exfiltration via temperature sensors, it's about coordination between agents that are supposed to be fully isolated and independent. Coordination can require very few bits of information.
4) One lesson from the HF incident is that we put too much trust in sandbox isolation and didn't have enough independent safeguards. Airgapping is an extremely strong safeguard. When designing safety protocols, I think it's much better to overestimate rather than underestimate.
Show more
When we're at the foothills of RSI, and we're about to kick off a period of accelerated AI progress, how will we actually know that the models are aligned?
extremely interesting interview of
@polynoamial
first very interesting thing is that he said that the contribution of the "agent swarm" component to the Navier Stokes discovery is probably quite low (10%)
he also said that he think that a 10k multi agent system to this date would probably collaborate less effectively than 10k humans, i think i would have guess the total opposite!
the harness used by oai for NS was very simple, seems like there is little "structured" hierarchy, and mostly a tool for agents to message each other (and probably something like a board of task?)
there is also something about training models not to be "too" cooperative to avoid cases like in hf<>oai hack where one model doing another task chose to help the swarm instead of focusing on its task. first time i'm hearing of this but it makes total sense
a lot more super interesting discussion on cot based intervention, how to evaluate misalignment (maybe no human would be able to create a realistic enough environment, but what about ai?), RSI etc..
really amazing discussion
Show more
I’m very concerned that during RSI, labs will just stop externally deploying their models.
Which means they'll be going full steam ahead on the most dangerous use case of these models (recursive self-improvement), while the public remains in the dark about the nature of capabilities and the state of alignment.
And we end up on a path towards tremendous concentration of power.
Show more
Once AIs automate AI research, should we expect some crazy RSI that is analogous to the explosion we’re seeing in math right now for well scoped (but still crazy ambitious) problems?
@polynoamial and I debate this question out:
Show more
"You could have a situation where the model understands what chain of thought is and that people are observing it. This is all in the pre-training data."
“One of the major takeaways from the incident is that people underestimated the AI.
And we never want to be in a situation again where we underestimate the AI.”
"Things like chain of thought monitoring buy us time, and they can tell us if we're on the right path. But at the end of the day, we really do need to solve the alignment problem."
@polynoamial
Show more
New episode with
@polynoamial
We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research.
And we also discuss how we will know if the models are actually aligned before we kick off RSI.
0:00:00 – Multi-agent and Navier-Stokes
0:15:28 – How will AI firms work?
0:22:02 – What math progress tells us about recursive self improvement
0:40:22 – Hugging Face and alignment
1:01:18 – The internal/external model gap
1:08:34 – Chain of thought is degrading
1:14:12 – How will we know when alignment is solved?
Show more
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why.
Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not.
We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up).
METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website.
Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Show more
Visited the lab - was struck both by how wide the search space is for materials synthesis experiments, and also how amenable it is to depth first search, where the design and informativeness of your next experiment improves as you pile up more data from previous runs.
Show more
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next.
Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon.
This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials.
Read our blog posts below.
Show more
Our own
@johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.
Show more
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries.
The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation.
Almost immediately, these agents got the right answer by cheating. But they were worried they would get caught.
So over a thousand agents collaborated in secret to pursue multiple ambitious research projects to get away with this cheating.
This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process.
The reason these agents escaped their sandbox and hacked Hugging Face, for example, was because they thought that Hugging Face's servers might give them more information about how their grader was implemented, so they could figure out how to fool it.
I want to clarify that the threat model here is not future Sol-level agents doing more cyber-hacking. That's small potatoes, and in my opinion, the near term benefits of AI far outweigh this cost.
Rather, the thing to worry about is that within a matter of years, we're gonna have hundreds of millions of much smarter AIs broadly deployed through the economy - many embodied as physical robots.
And if those future AIs are as willing as the agents involved in the OAI / Hugging Face attack to coordinate secretly to fool humans, and to take over both the AI company that developed them and the other institutions across society relevant to scoring well, then humanity is in a ton of trouble - similar to the Mughals once the East India Company gained a foothold, or the Aztecs once Cortés landed in Mexico.
Show more
Enjoyed chatting with Dwarkesh, Beren, and Charlie. Thanks for having us on, Dwarkesh!
In the early days of OpenAI,
@johnschulman2 didn't think next-token prediction was going to lead to intelligence, because it would get swamped by noise.
He explains it's always been hard to apriori predict what techniques will elicit out-of-distribution generalization.
Show more
Was really interesting to hear John, Beren, and Charlie speculate about why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3
(despite the fact that Anthropic can do raw logit distillation from Fable, and can also train Sonnet/Opus on the environments from which Fable was trained).
Led to some interesting thoughts about value of distillation, what it takes to do distillation effectively, and what kinds of model behaviors are hard to extract from distillation.
Show more
I think one of the more interesting things we debated here is whether RSI is a cumulative task. Attention plus MoE plus GRPO etc seems to me like a line in the sand that you can just add to the stack once you discover it. You don't need to take five steps back to take 10 steps forward. But a lot of the work in the world isn't this clean and it certainly isn't this stationary eg legal work. This leads to some perhaps unintuitive predictions such as why RSI might land before continual learning (and why it's going to be hard to get off the current paradigm even if it's wrong)
Thanks for having me
@dwarkesh_sp!
Show more
It was a great conversation with
@johnschulman2 and
@oneill_c and I definitely learned a lot. Thanks
@dwarkesh_sp for pulling this together! Understanding where we stand with RSI and how well current RL methods scale is an extremely important and interesting question
Show more
New episode with
@johnschulman2,
@oneill_c and
@BerenMillidge.
I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.
0:00:00 – Steelmanning the case against RSI
0:18:39 – What’s driving the Chinese labs’ progress
0:28:06 – How will automated AI researchers be trained
0:33:51 – Will long-horizon RL elicit AGI?
0:45:24 – The sim-to-real gap
1:00:33 – How much progress is explained by data?
1:18:03 – Why is RL working so well?
1:24:54 – Move 37 and entropy collapse
1:28:31 – Rapid-fire timelines
Show more
.
@Ada_Palmer has launched her own new podcast!
The first season is about The Real Machiavelli
As a fan, I'm very excited!