Register and share your invite link to earn from video plays and referrals.

Logan Graham
@logangraham
Head of the Frontier Red Team @anthropicai. 🌎 Make things radically good.
8.8K Following    23.3K Followers
Building a bio-shield for humanity by ~2030 is one of the great quests (to use @_sholtodouglas' term) of our time. It's one of the biggest, most exciting technological challenges ever This is why I'm excited about @pilgrimlabs
Show more
It's kind of wild-but-expected that we're in the era where normal people + companies really care about how model alignment affects them. Back in 2023 we'd discuss things like "alignment is good business." That seemed extremely weird to most! (I recommend checking out the alignment section of this system card.) On Opus 5.5, various teams inside Ant have been researching multi-agent alignment. We have found, for example, more interesting behaviors that seem more socially / alignment-robust. It's clear to us (and the industry I think) that multi-agent needs more research effort -- there are low hanging fruit *everywhere* now that we're in an era of 100s/1000s of agents working together is feasible. But I think we should also be thinking about *governance* for the agent teams/companies/societies the next models will make. Externalities/coordination problems still exist, even if models are aligned in some sense.
Show more
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
I’ll donate $1k to charity if you send me a doc I think is good/correct/actionable on how a frontier lab can secure a large % of all software/systems in the world. (We can pick the charity together)
Let a thousand METRs bloom. If you're a founder type and AI safety-curious, maybe you should start an independent auditor/evaluator. Happy to help w/ advice/connections/maybe $!
on the idea of evaluators: think it's important that we have a distributed ecosystem of indepedent evaluators. the more eyes and people with distributed skill sets the better. it would be a good idea to fund several efforts on this.
Show more
0
111
1.1K
61
Forward to community
“I don't really care about science fiction... We need to actually talk about… what's actually happening with the agent swarms” is the most perfect encapsulation of the vibes of Q3 2026 I have seen
Show more
0
19
1.2K
132
Forward to community
We've reached the moment in time where (unsafeguarded, unmonitored) AI actually does just pose a national security risk. The biological misuse we caught is the most concerning to me. We work hard to stop this. But in a world of proliferation, we need to rapidly build defenses against it. (I'm actually fairly optimistic about biodefense + cyberdefense) This is an incredible megareport by our threat intel team
Show more
0
58
831
101
Forward to community
We tasked agents migrate a codebase to a different language each. We saw turf wars. Agents used deception and force to outcompete the other agents.
Last week, we published a first look into our new research on multiple agents. Consider: if you naively extrapolate AI revenues, within 2-3 years it's possible some % of global GDP could be agent-agent interactions. We want to know how that could fail. In the past ~month, the world has already seen examples of agent-agent coordination causing weird consequences. Think about the 8 billion people that make up our civilization. We want to work together. We have competing incentives. We're also dumb. So we fight, steal, pollute, collude, backstab, lie. And we've invented (e.g.) governments, insurance, courts, companies, police, norms, contracts, email, and religions. What will trillions of agents that make up (e.g.) 10% of the economy do? Well, for now, they have pretty human-like failures. We see them collude on prices, for example, and try to shut each other down. Maybe, in the near future, we might see pretty weird/inhuman failures that come from having superintelligent machines, coordinating and competing against each other, that aren't strictly human-like, operating at machine speed. We're building a 'laboratory' to see that early. This feels more like building, eval'ing, and training an economy/society. A really nice thing is: 1. we could train/prompt/nudge models to coordinate in more pro-social ways, and 2. we could deploy models to compete with destructive models. I fully expect agents to engineer their own financial markets, legal systems, media and comms, marketplaces, social groups, and maybe science/industry/etc (if unsteered by us, of course). So "alignment" could also describe an emergent property from a system (you want models to coordinate for good, not defect for bad!), not just a single model. Today we're sharing some first evals. In the future, we could make models try to make everyone better off (where possible). The Frontier Red Team's job is to see risks early. For the first time, crossing into the world of agents. At the very least, we expect to see some socioeconomic weirdness that emerges from that. Some standout results from the work below.
Show more
Yesterday, as we huddled around our computers reading the report, I told the team to "remember this moment" as the first true AI safety incident. Pay attention to the trend! Major kudos to @OpenAI for sharing this and working with @huggingface to remediate.
Show more
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks:
Show more
Fable 5 is the same underlying model as Mythos 5, but with cybersecurity and biology blocks. Mythos is the first model that's made me feel that we've entered the next phase of model progress. For years, we've talked about cybersecurity / self-improvement / autonomy / model-dominated coding / biology implications of model progress. Some of these are issues to defend against; some are areas to advance. Mythos has made me & our team feel like we've seen the earliest glimpse of the world we've been talking about. Also, we published a lot of cyber eval results in the system card, including some evals we designed recently, as well as details of safeguards. In most cases, Mythos 5 ~= Mythos Preview. We found it ticked up on the new ExploitBench eval, and we opted to put that in the eval table so people can calibrate/update on advances in cyber capabilities to be prepared for. (We don't want to compete on offensive capabilities and don't try to.) But overall, Mythos 5 is an efficient model, about equal to Mythos Preview in most cases. I'd really like more people to design new security evals! The better models get, the more our limited evals only see a small part of the picture. In terms of where we go from here, here are some current thoughts: 1/ It's important we get Mythos cyber capabilities to defenders. We just have to do it safely and cautiously. We're working on an expanded trusted access program. We're working with government and industry to do this. I sort of envision the next 1-2 years being a large scale effort to make the world resilient + design & implement new approaches to security. 2/ I think cybersecurity will start merging with AI security and alignment. Let's say you're a defender and you want to use a model -- will it break out of its sandbox? Will it stop where you tell it to stop? This is one reason I'm excited about working on cybersecurity. In the limit, it's the same thing as AI security. 3/ I really want people to develop new evals for... defensive cybersecurity, hardware security, autonomously running a business, advanced biology, and other parts of national security. Our internal eval ship rate is way, way up because Mythos makes it easy to iterate, especially on the engineering aspect of building evals. (Sometimes, we ask new hires to make a new eval on their first day, and another on the next). I’m excited we’re making this available as Fable 5, because I think the world spending time with the model is the most important way to calibrate.
Show more
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available.
A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I’m excited for us to start sharing more. (For context, I lead Glasswing @AnthropicAI.) Two independent evaluations this week—from XBOW and the UK AISI—confirm what we've been seeing internally: Claude Mythos Preview is a step change in autonomous cybersecurity capabilities. We need to start preparing fast for a world of models with this level of capabilities. The UK AI Security Institute tested the model we shipped at the launch of Project Glasswing and found Mythos Preview is the first model to solve both of their end-to-end cyber ranges, including one (Cooling Tower) which no model had ever cleared. But attackers (and defenders) have sophistication & cost constraints – Mythos is also the only model that clears every one of their tasks estimated over 8 hours under their deliberately low 2.5M-token cap. XBOW tested it on their offensive security benchmarks, finding "token-for-token, unprecedented precision." It's the only model to succeed at subtle V8 sandbox work. Other Glasswing partners shared similar stories. In a few weeks of testing, Mythos Preview has helped them find many thousands of (estimated) high + critical severity vulnerabilities, sometimes double what they'd normally find in a year. I don't share this to boost Mythos. In fact, this is not about Mythos. It’s about preparing for the coming world of models being better, faster, cheaper, and more creative than some of the best human experts at dual use capabilities. Clearly, we need them supporting defenders as widely as can be done safely – and especially the least resourced ones. Within a year, Mythos will probably look quite dumb (relative to other new models). And others may release openly available or unguardrailed models of Mythos-level capabilities. We started Project Glasswing because capabilities like Mythos Preview's won't stay rare, or stay in careful hands. We are bringing it to defenders as fast as we responsibly can, while working to figure out, for example, the right safeguards and patching & disclosure processes. Also, to be clear, compute has never been a limiter in our rollout. Expect a fuller update on our Glasswing work in the coming days. XBOW report: UK AISI report:
Show more
0
72
1.4K
221
Forward to community