Register and share your invite link to earn from video plays and referrals.

Dan Altman
@manaltdan
AI Policy @AnthropicAI, focusing on model development & research. Previously AI governance & geopolitics research @GoogleDeepMind
1.5K Following    1.4K Followers
We released some new ideas and data on how to measure: 1) the extent to which AIs are being used to build next-gen AI models, 2) how we exercise oversight over increasingly autonomous agents, and 3) how compute is allocated. Measurements like these will be essential to future AI policy regimes. Enjoyed contributing to this work!
Show more
the day pangram integrates into streeteasy it’s so over for these glorified hellholes
1/ Some personal news: I'll be joining @AnthropicAI as Head of Frontier Compute Strategy to work with @NotTomBrown on our infrastructure expansion, coalition building, and planning for rapid AI progress.
Show more
0
137
1.9K
77
Forward to community
METR is just one of many existing AI evaluators! If you take issue with METR, that isn’t a good argument against requiring companies to carry out embedded audits Here are 20 other evaluators who work with AI labs to assess risk: Transluce Grey Swan Apollo AVERI SecureBio Faculty Vaultis Dreadnode Irregular RAND Far AI ActiveFence (now Alice) Active Site Deloitte (via Gryphon acquisition) Nemesys Mercor (via Sepal acquisition) AE Studio Scale Frontier Design Redwood
Show more
I’m really excited and optimistic about this. People disagree on nearly every major crux of AI development, but employee-like access for evaluators is upstream on what almost all of them want, whether it’s generally better transparency, more info on misalignment, incident reporting, bio risk assessments, pacing. The list goes on.
Show more
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
Show more
Anthropic is unilaterially committing to embedding evaluators to verify safety practices and report incidents. Evaluators will get access badges, company laptops, and "permissions and tools similar to those of internal employees who do comparable risk assessments." This is a remarkable step for frontier AI auditing, transparency, safety, and security. I hope that other companies follow suit. And I look forward to continuing to build the standards, practices, tools, and policy at @AVERIorg to make this work rigorous, effective, and universal.
Show more
"The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents."
Show more
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here:
Show more
Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same.
Show more
0
235
6.4K
702
Forward to community
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
Show more
0
297
6.6K
1.1K
Forward to community
I love Foreign Affairs but get a good chuckle every time a new cover goes up and it's a remix of one of: the world order is fracturing, China is rising, tech is geopolitics now, or America's alliances are crumbling
Show more
The September/October 2026 issue of @ForeignAffairs is now available. The latest issue provides analysis on the future of U.S. alliances, the threat of a global economic crisis from Chinese overcapacity, the dangers of synthetic pathogens created with AI, and more. Start reading:
Show more
It really is almost unfair how beautiful SF is
San Francisco is the most beautiful city in America by far, and it's one of the most beautiful cities I've ever seen.
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
Show more
0
939
18.3K
2K
Forward to community
A defining feature of mid-2020s product design seems to be placing LLM features in the places you are most likely to accidentally press them. One day I will be nostalgic for it
Me sleeping peacefully at 11:59pm while my AI agent prepares to secure a 6pm court time at NYC Riverside Park.
the sf tennis reservation system will become one of the most hardened softwares on the planet of earth
0
18
4.1K
112
Forward to community
Powerful AI will change our institutions. Yet most of our attention is on aligning or regulating AI, while treating our institutions as fixed. Introducing Pax Machina: a new publication about the institutions we need for powerful AI. It is edited by @MaxKronerDale, @sebkrier, @NoemiDreksler, @LiamPatell, @chelcott9, @synchroaphasia, @ryan_t_lowe, @edelwax, and @klingefjord. The editorial board consists of Peter Railton, @sethlazar, @saffronhuang, @AmmannNora, @xuanalogue, @IasonGabriel, @hamandcheese, and @deanwball. Our goal is to seed rigorous debate about what a world of humans and powerful AI could and should look like. We will publish proposals for new institutions, counterproposals, well-justified design principles, and analyses that change how a class of institutions should be designed.
Show more
0
61
1.1K
190
Forward to community
Briefly daydreamed about one of my future descendants reading the wikipedia entry for endangered LLM vernaculars. They scroll to a subsection, Revitalization Efforts, about a recent movement of cultural preservation agents reseeding the corpus with "absolutely!," the em dash, "one honest caveat:"
Show more