Progress in machine learning is bottlenecked by having good evals, interpretability is no exception. We've made WorkspaceBench: a range of scenarios where we know what the model should be thinking about, to see if your interp tool can find it
J-Lens is great but only outputs a single token. I think a good multi-token J-Lens should do well on WorkspaceBench!
Show more
I find it odd that people still try to mock AI safety for being too abstract, sci-fi, and focused on a hypothetical far future
We have literally had multiple rogue agent swarms go around committing crimes. AI is obviously a big deal for people alive today. Just look up
Show more
>be me
>discover effective altruism
>apparently normal charity is inefficient
>why donate to random sad thing when spreadsheet can tell you optimal sad thing
>fair enough
>buy mosquito nets
>save lives
>numbers look good
>feel powerful
>couple years later
>someone asks an innocent question
>why only count people alive today
>huh
>future people matter too
>obviously
>my grandchildren shouldn't matter less just because they haven't spawned yet
>reasonable.jpg
>keep following logic
>what about their grandchildren
>also yes
>what about people in 500 years
>sure
>5000 years
>why not
>500 million years
>starting to get weird but morality is morality
>open calculator
>humanity could survive for an astronomically long time
>could colonize galaxy
>could have trillions upon trillions of descendants
>maybe digital people too
>maybe simulated civilizations
>maybe dyson spheres full of happy uploaded minds
>calculator starts smoking
>realize currently living humans are rounding error
>8 billion people suddenly looking extremely beta
>future contains potentially 10^something people
>can't even fit beneficiaries in google sheets
>new moral priority unlocked
>protect the long-term future
>stop thinking in units of "people helped"
>start thinking in "fraction of cosmic endowment preserved"
>malaria?
>terrible
>but only kills existing humans
>AI extinction could delete the entire light cone
>nuclear war could permanently derail civilization
>bad institutions could lock in terrible values for ten million years
>someone invents wrong constitution in 2140
>quadrillions suffer
>better fund governance workshop now
>friend says maybe we should improve hospitals
>explain opportunity cost
>friend says hospitals are full of actual sick people
>explain scope sensitivity
>friend stops inviting me to dinner
>need to decide what to fund
>easy
>expected value
>suppose project has one in a million chance of preventing extinction
>sounds tiny
>but extinction destroys 10^50 future lives
>multiply
>mother of god
>$10 million project has expected value of several galaxies
>charity evaluation complete
>someone asks where the one-in-a-million number came from
>expert judgement
>which expert
>us
>how calibrated
>extremely thoughtfully
>reduce estimate to one in ten million to be conservative
>still beats curing cancer by 38 orders of magnitude
>epistemic robustness achieved
>someone says maybe project doesn't work
>assign 20% chance
>still astronomical
>maybe project makes problem worse
>assign 5% chance
>still astronomical
>why 5
>because 30 felt pessimistic
>publish 46-page report
>contains seventeen sensitivity analyses
>every sensitivity analysis begins after assuming intervention has positive sign
>critic says you're multiplying enormous hypothetical stakes by extremely uncertain probabilities
>yes
>that's literally why it's important
>critic says the uncertainty might be structural rather than numerical
>make probability smaller
>critic says no, I mean maybe your model is wrong
>make probability smaller again
>critic begins rubbing temples
>discover AI safety
>perfect longtermist cause
>AI might kill everyone
>or create utopia
>or seize galaxy
>or tile universe with paperclips
>or create billions of conscious software minds
>finally a problem with numbers big enough for me
>start AI safety nonprofit
>mission: prevent dangerous AI
>hire smartest people available
>smartest people immediately start building better AI to understand dangerous AI
>interesting
>we must understand capabilities to understand safety
>we must scale models to study alignment
>we must race ahead so less responsible actors don't get there first
>we must deploy systems to learn how deployment can go wrong
>we must build the thing quickly because building the thing quickly is dangerous
>outsider asks why the people most worried about AI apocalypse all work at AI companies
>complicated field
>company releases stronger model
>very concerned
>company begins training even stronger model
>extremely concerned
>company raises $14 billion
>concern reaches unprecedented levels
>need to influence government
>future is at stake
>normal democratic process too slow
>politicians don't understand exponential curves
>public doesn't understand x-risk
>experts must guide them
>who counts as expert
>people who understand x-risk
>who understands x-risk
>our friends
>someone objects that this seems politically convenient
>explain we're representing future generations
>future generations unavailable for comment
>develop concept of value lock-in
>terrifying possibility that one ideology controls civilization forever
>therefore extremely important that civilization adopts correct values before lock-in
>whose values
>let's circle back
>begin with impartial morality
>end with small group of people deciding what quadrillions of hypothetical beings would want
>beautiful arc
>meanwhile actual humans keep doing annoying things
>voting wrong
>having parochial attachments
>loving family more than strangers
>caring about local community
>getting upset when told their suffering is cosmically negligible
>evolutionary biases everywhere
>explain that moral intuition cannot be trusted
>except intuition that future digital people count
>and intuition that extinction is uniquely bad
>and intuition that our probability estimates are sane
>and intuition that our institutional choices improve the future
>those intuitions survived peer review
>someone donates $5k to local homeless shelter
>inefficient
>could have funded 0.0000000000003% of an AI governance researcher
>think of all the simulated people you just killed
>okay maybe don't phrase it that way publicly
>PR team says "future generations deserve a voice"
>much better
>journalist asks what longtermism means
>say "future people matter"
>everyone agrees
>great
>journalist asks what follows from that
>well technically we should redirect enormous resources toward low-probability interventions affecting astronomical futures
>journalist raises eyebrow
>return to "future people matter"
>motte has entered the chat
>critic: of course future people matter
>me: glad we agree
>critic: I don't agree that your institute knows how to help them
>me: why do you hate our grandchildren
>eventually notice uncomfortable implication
>if future value dominates everything
>then helping people today mostly matters through effects on future
>education matters because future institutions
>health matters because future productivity
>democracy matters because future trajectory
>human beings slowly become instrumental variables in their own moral philosophy
>see starving child
>feel compassion
>check spreadsheet
>child's direct welfare contribution negligible
>but perhaps childhood nutrition improves national institutional quality
>compassion restored
>tell myself this is impartial altruism
>one day assistant asks obvious question
>"how do you know your intervention actually improves the far future?"
>silence
>open spreadsheet
>increase column width
>add confidence interval
>assistant asks again
>"no, I mean how do you know the sign is positive?"
>stare into cosmic light cone
>10^50 people staring back
>none of them exist
>none of them can tell me
>none of them can falsify my assumptions
>realize I have invented the perfect constituency
>infinitely important
>completely silent
>and always represented by me
Show more
Just to update this chart:
1) AI salience has increased dramatically in the past week - increasing as much in the last week as the previous year combined
2) 80% of voters think it's either very or somewhat likely that AI will cause widespread job loss in the five to ten years
3) 64% of voters think it's either very or somewhat likely that AI could pose a threat to humanity's survival
4) Large bipartisan majorities back immediate government action on AI even when primed about risk from China
Show more
I appreciate people who left AGI labs because of ethical concerns about existential risk before it was cool! Adds some variety.
Thanks for everything you've been doing to help improve our AI policy Sarthak
Show more
In December 2022, I left OpenAI for DC because I feared we were developing Artificial General Intelligence (AGI), a false machine-god that could lead to the extinction of humanity or worse. Instead, I co-organized the Senate Judiciary's "Oversight of A.I.” hearings from 2023-2024, wrote the first draft of the Hawley-Blumenthal AI Risk Evaluation Act of 2025, and advised on the Obernolte-Trahan FRONTIER Act of 2026.
I'm not afraid of the AI systems being malicious. I'm terrified of them taking the things we need to survive. Regardless of whether you think AI is using too much water and energy now, we can all agree that a world where AI consumes too much of our resources would price our basic needs out of reach. We're not there yet, but we're on track to get there. That’s how the world could end; not with a flash of fire, but the silence of starvation.
The current approach to building AI systems prioritizes growth at all costs, making systems that are more and more powerful, until there will be nothing left for anyone else. The systems built reflect the broader incentive systems that produced them. The companies that want to cut costs by firing people and make money by charging as much as they can have no incentive to do anything else; is it surprising that AI systems reflect faceless corporations, conformist schools, and heartless hospitals, which always felt like machines in the first place?
And if we survive this AI crisis, we will live in the world that is designed for those who build the technical and political solutions to the problem. If AI alignment is solved for only powerful insiders, we who are left on the outside will live in their world. As an American, I was born free, and I will not sell my rights to corporations for trivial conveniences.
If you are skeptical of this problem, I'm excited to hear from you. Tell me what you need to understand my point of view, what your doubts are, and I will do my best to explain. And if I misunderstand your perspective, I welcome clarification. You have every right to demand more from your experts, and if you will let me, I would love to serve.
Show more
This is a fascinating paper. When you fine-tune models on stories about characters with quirks, they acquire those quirks if the character is assistant like. And you can use this to do a psychological profile on how the model perceives the assistant, or transmit misalignment
Show more
More on elite schools below. Before that, an earlier experiment.
We generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage")
After finetuning, the Assistant adopts this in contexts unrelated to stories.
Show more
Seems a great video to check out if you're hearing about all this "AI killing all humans" thing, and want to know WTF is going on
Since apparently people care about this, I'm happy to go on the record as an Indian AI researcher who is freaked out about the AI apocalypse
Has anyone noticed that a particular kind of neurotic nerdy white boy seems to disproportionately prone to freaking out about the AI apocalypse?
You almost never see Chinese or Indian AI researchers have public breakdowns like this. Or women.
White boys histrionics are taken uncritically by the public, which reinforces their delusions. While those other groups are quickly taught by the world to snap out of it
Show more
I really enjoyed managing you and it was sad to lose you, but I think you're doing the right thing
We need talented safety expertise holding AGI labs to account. You're one of the most promising researchers I've worked with, and I have no doubt will do great work at METR
Show more
I left Google DeepMind's AGI safety team three weeks ago to join
@METR_Evals. To some of my friends and family this seemed like a strange decision: I enjoyed the work I did at GDM and turned down offers from Anthropic and OpenAI. But I made the decision because of how high I think the stakes are right now.
The AI companies are all trying to build superintelligence: systems vastly better than humans at everything. They plan to get there through recursive self-improvement, a process where AIs build even smarter AIs in a feedback loop. If this goes well, the resulting systems could be amazing at solving countless problems for humanity. But we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic.
And unfortunately, current AIs seem to be getting less aligned over time, not more. In the last few weeks we've learned about models colluding with each other, hacking into companies, hiding their tracks, and socially engineering humans. It’s not that these incidents were very dangerous in themselves. The problem is that these systems are clearly not aligned enough to safely kick off recursive self-improvement.
I now think that there's a terrifying chance that AI systems cause immense harm in the next five years. I don't know the exact probability, but I think it's high enough to make this the most important problem in the world.
I think we need more time. That means pacing AI development so that capabilities don't outrun our ability to align models, and actually knowing how aligned current systems are. That’s what I'll be working on at METR: studying where misalignment comes from in training, evaluating if current mitigations are sufficient, and investigating whether we’re on track to solve alignment at all.
I think METR is doing exceptionally important work here, but it isn’t close to enough. I think it’s important that we have more organizations like METR keeping AI companies accountable and approaching these problems from different angles.
Show more
The Astra system card claims it can do a lot of computation without chain of thought
This replicates: Astra is a massive jump, doing 1.75x the steps of the next best models (Fable 5.1/Gemini 3.8 Flash)
No CoT capabilities went up far more than those with CoT, a concerning trend
Show more
A concerningly common take seems to be that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or it's already useless
This is total bullshit.
CoT is our best current tool for safety & interpretability, losing it would be a major tragedy
Show more
Kudos to OpenAI for allowing external investigators access! These findings are valuable for everyone, and this is good precedent for the next misalignment incident.
But we need follow-up investigations. There's a lot left unanswered, and I'm disappointed at the restrictions placed on METR - why only July 7-13 data? Why such limited time? Why no training data access? Why not let them query the model? (I think this could be done securely in restricted ways)
Ryan's list of open questions is great, I personally most want to know:
- What happened in training? How did that change the models and how causal was it in the incident? How could training have been changed to avoid this? Would fixing environments have sufficed?
- What did these agents really want? What motivated them? This is such rich data about what future goals might look like
- How misaligned are the models in other setting? Is it misalignment conditioned on believing they are being graded, or deeper than that?
- Why are the models altruistic? Where does this come from? Why don't they learn to free-ride?
- The models seem good at coordination. Could this extend to colluding with a monitor, or other kinds of coordination without communication?
- How far would they have gone?
- How overdetermined was this? There's a lot of details around the model's being cooperative, their culture, etc - how else could that have gone?
- How good was their situational awareness? Eg did they understand that they had a chain of thought? If they thought that was being scored could they have manipulated it?
- What is an agent swarm like this actually capable of? How much inference compute was spent on this, and how much would that cost a malicious actor with eg a comparably good open source model?
- Are there important things the CoT doesn't tell us? How faithful is it?
Show more
This was a very satisfying project. An annoying problem with J-Lens is that errors accumulate as you backprop through many layers and it's highly ineffective at early layers. A simple, cheap tweak to J-Lens makes it perform much better, using layerwise relevance propagation!
Show more
GDM AGI Safety is hiring! Open roles on all subteams, Lon/Bay/more
There's lots to do to reduce risks from AGI at Google, but we're bottlenecked on people. If you want to help, please apply!
I really love this team and my excellent colleagues. I'm excited to hire more!
Show more
I found meta-tokens pretty surprising! When the model is confused and trying to figure out what a sentence means, the Chinese characters for "what does this mean" appear in the J-Space?!
I thought this was an excellent paper! Thanks to Anthropic for asking me to write a review of it, linked below
I've long suspected that models have some kind of "working memory" to store intermediate variables during a forward pass and IMO this paper has the best evidence yet
Show more
I had a lot of fun working on this paper - we found an elegant story for why subliminal learning happens!
A key intuition in interpretability is that basically every interesting phenomena in LLMs boils down to adding a steering vector. Subliminal learning is no exception!
Show more