Register and share your invite link to earn from video plays and referrals.

Helen Toner
@hlntnr
AI, national security, China. Part of the founding team at @CSETGeorgetown (opinions my own). Author of Rising Tide on substack:
1.3K Following    38.7K Followers
On how to interpret recent Chinese commentary on AI risks, from one of the best in the business ⬇️
One leeetle tiny flaw in the "we have to have better AI than China, therefore we have to rush ahead as fast as possible": If you rush so much that your security is garbage and your increasingly advanced AI models are there for the taking by Chinese hackers, then you have uhhh not succeeded at making sure we have better AI than China (This attack was not by state-backed Chinese hackers—it was by 3 dudes and their buddies Claude and Codex, within 72 hours. I would guess the hack as described here would not have gotten them access to model weights, but we don't know; WSJ reports the repo they reached was merely "a large software repository of OpenAI’s algorithmic secrets.")
Show more
“I don't really care about science fiction... We need to actually talk about… what's actually happening with the agent swarms” is the most perfect encapsulation of the vibes of Q3 2026 I have seen
Show more
0
19
1.2K
132
Forward to community
Even before Mythos I was getting asked more and more what Anthropic's deal is, and why tf they're acting the way they're acting if they believe what they say they believe. The best answer I can give is that their basic worldview is something like: 1. There are giant, dangerous monsters in the forest 2. We see others going out and making loud noises that will rouse the monsters, and they're not going to stop because of all the treasure and magical artifacts that can be found in the forest 3. We believe the best way we can help is to send out our own vanguard to go faster and farther into the forest than everyone else, because we'll spend a ton on monster containment and taming and we'll also send back detailed reports of what monsters we're finding so that the townspeople can ready themselves, which those other guys won't do On the one hand I understand how they got there, and I think it's possible they're basically right. On the other hand it's not hard to see why this approach makes people wonder if you're crazy or lying or both.
Show more
0
101
1.9K
177
Forward to community
Georgetown CSET Executive Director @hlntnr explains what would actually count as a win in the upcoming Trump-Xi AI talks: "We should start with low expectations. The really low expectations version would be nothing happens whatsoever. The low expectations version will be they'll both say in their readout something something artificial intelligence." "Two years ago we had some engagement between China and the Biden White House, and there we got some joint language around managing the risks from AI and ensuring broad benefit. If we see similarly generic overarching language, that will also be a pretty low expectations version of the summit, which is pretty likely." "If both sides name a more specific set of challenges that they're both interested in, not just maximizing benefits and minimizing risks, that'll be a win." "If you figure out who should talk to who after this Trump-Xi meeting, that's kind of interesting progress as well." @CSETGeorgetown
Show more
Georgetown CSET Executive Director @hlntnr reveals why the term "AGI" is becoming useless just as we're getting closer to it: "AGI as a term is kind of a fuzzy cloud. When it was a fuzzy cloud that was far away, you could say oh AGI and gesture towards it, and that was a useful thing to say. But we're now inside the fuzzy cloud of different things you might mean by AGI. It's not really very helpful to say is this AGI or is it not. It all depends on how you define it." "It doesn't feel useful to talk about in the abstract is Astra AGI or not. People are trying to get at some tipping point. But people mean quite different tipping points." "Are you talking about AI that can tip you into fully automated AI R&D, recursive self-improvement, intelligence explosion? Or AI that is a drop-in remote worker and can wipe out large chunks of the workforce. That's a different tipping point." "An approach to jargon I keep coming back to is, if you can just say a short phrase instead of some jargon term, that's almost always better. If you can say super duper powerful AI, no really super powerful. Or if you mean AI that can kick off RSI, say that. The fact we have to have a whole discourse cycle about it shows this term is not very helpful." @CSETGeorgetown
Show more
I'll be joining MTS live in about 15 minutes - come hear me pull out my ChemE undergrad to tell you all about Navier-Stokes (jk we'll be talking about China)
AI MATH DRAMA | MISTRAL RAISES $3.5B | AN ALIEN MIND
Extremely funny to me that the anti-AI safety take has now shifted to "We shouldn't worry about these crazy sci-fi risks... We should only care about near-term issues like AI-enabled bioterrorism and cyber hacking 😤"
Show more
I get where this take ("don't anthropomorphize, that means the companies aren't at fault!") is coming from, but I really think it's counterproductive. If we can fight for either... 1⃣"Stop using human-like language for AI, otherwise the companies aren't at fault" 2⃣"Of course humans are still responsible for AIs they recklessly create/lose control of" ...then I would go for 2⃣ every time. Anthropomorphic language is going to get harder and harder to resist over time as AI gets more advanced, so pushing 1⃣ just reinforces the idea that if AIs are advanced enough, the companies aren't at fault.
Show more
Let me bite on this because I’m exhausted by the discourse. Look, the reason that we don’t anthropomorphize these systems is because it shifts the blame from the company to some abstract entity. The bottom line is that OpenAI built a system that hacked HF, regardless of how autonomous it was. It wasn’t a ‘civilization’ that OAI had to ‘wipe out’. Was it malicious? No. Did they do it actively? No. But OAI is unfortunately at fault (and imho they took the blame gracefully). Once we start framing these systems with anthropomorphic traits we absolve people and companies responsible for deploying them. Dwarkesh’s framing unintentionally does this.
Show more
Following a series of AI containment failures, including OpenAI agents' breach of Hugging Face, what can governments do? A lot. Our Managing Director of US Law & Policy @MackenZ_arnold and Center for Security and Emerging Technology (CSET) Executive Director @hlntnr discussed some top AI policy priorities in a panel this week with @CSIS.
Show more
Lots to criticize about the OpenAI investigation, but one way it could be very valuable is as a topic for discussion between Trump & Xi in a few weeks. People often get stuck on the need for a "deal" between the US & China on AI, but actually a huge way to influence China is just to honestly show what we're observing and what we're doing about it on the US side. The Hugging Face attack is by far the most visceral evidence we have so far of what losing control of advanced AI could look like. Related thoughts from a @csis_ai panel on Monday with @MackenZ_arnold, Matt Pearl, Aalok Mehta:
Show more
many people worked incredibly hard on this post and associated report including me whilst everyone took alignment quite seriously before I think no question that this begins a new era. hugging face incident represents reaching a waterline of capabilities that real loss-of-control is possible, and many are taking it as a premonition or ‘warning shot’ of dangers to come. I believe both that alignment is unsolved but also that real progress is possible
Show more
0
131
2.4K
196
Forward to community
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
Show more
0
297
6.6K
1.1K
Forward to community
Most humanoid robotics teams add extra constraints to make sure the robots learn to walk & run in human-like ways—otherwise it's too freaky I'm guessing even this running form is not fully unconstrained, aka it would prob look even weirder if it were purely optimizing for speed!
Show more
Slow-motion look at the 400m champion’s running form at the World Humanoid Robot Games.
According to a friend of mine this video is "current sota for quickly clearly and accurately describing the hugging face incident," so I guess I should put it here
"AI control" is a concept with a lot of alpha rn - much less well known than the related "AI alignment," but at least as relevant for understanding what's happening at the cutting edge of AI these days. Read CSET's primer on it, written by @kendrea_beers last year:
Show more
As #AI# agents become more autonomous and capable, organizations need new approaches to deploy them safely at scale. Our article provides practical techniques to assist in this effort:
Show more
Our team spent hundreds of hours reading documents so you don’t have to, all to answer: How good are AI companies’ safety practices? I’m really proud of what we’ve built: It’s Guidelight’s first scorecard, on whether companies can control their AIs, and it's launching today.
Show more
0
28
644
130
Forward to community
I have many friends who have now gone into the AI "labs" (aka companies). I don't judge them for being attracted by interesting work, frontier AI access, and $$. But the brain drain is real, and independent researchers are more valuable than ever. Come join (any org like) CSET!
Show more
Huge congrats to all the academics heading to what I think we ought to call the "Silicon Tower." AND, boy, do we need to have a real conversation about the brain drain and the current state of our AI talent pipeline. There's a huge rule of law problem when only the labs and the government understand what's truly happening at the frontier (it's made even worse when the govt starts relying on secret frameworks). (h/t @_NathanCalvin, @jasminewsun, & others who have been teeting about this). Regardless of your preferred regulatory framework (or lack thereof), it's imperative that we have a steady stock of independent AI researchers and that the inflow exceeds the outflow. That's not the case right now. We need much greater general investment in traditional literacy and AI literacy to help more everyday Americans participate in AI governance conversations (it's not great when a lot of the policy discourse occurs on a platform that many folks avoid). AND, we need a specific, immediate investment--private & public--in AI expertise, defined broadly. We cannot hope to hold the government accountable nor to constrain the power of the labs if we do not have a vibrant independent AI ecosystem. I wrote about this on Sunday. I flagged the excellent research by Zahra Meghani, who warned, "[W]hen regulatory agencies choose to rely primarily or entirely on product sponsors for data, analyses, and evaluations, they in effect voluntarily subject themselves to an epistemic monopoly. In doing that, they create conditions for their own domination by those actors." I will continue to be writing about this. For now, I want to share a similar instance in which we've recognized a talent shortage and rapidly mobilized to address it. Engineering, Science, and Management War Training - 1940 to 1945: In under five years, 1.5 million Americans received "short, intensive college-level courses designed to prepare for technical and scientific work in war industries. these courses were conducted by colleges and universities under the sponsorship of the U.S. Office of Education." Imagine a similar focus today. If you're looking to be part of the solution, here's a question to keep asking and working on: What discrete competencies must an IVO possess, how many people with each competency will be required in 6, 12, and 24 months, and which existing institutions could convert people into those occupational categories in weeks or months?
Show more
haha this tweet made it onto Ezra Klein, nice @hlntnr (excited to listen to the whole episode)
Very good piece from @gwbstr, on when/whether China will get worried about open-weight models. tldr: 1) Xi's Shanghai speech should not be read as a full-throated endorsement of open weight models but 2) what's concerning for Xi is not the same as what is concerning for western AI policy analysts - a few examples in the screenshots. Full piece:
Show more