Register and share your invite link to earn from video plays and referrals.

Christian Catalini
@ccatalini
Founder w/roots in academia. Founder @MIT Cryptoeconomics Lab. Past: Co-Founder & Chief Strategy Officer, Lightspark. Co-Creator, Libra. Head Economist, Meta.
5.8K Following    26.2K Followers
With the external hack of OpenAI via Claude, closed models continue to be the tip of the iceberg on AI risks, not open models. They have been 1) easier to get started with, 2) more capable & 3) shipped w/ leaky safeguards. Finetuning open models to specific attacks is harder.
Show more
The paper he is referring to is BitWhisper from 2015. The cpus had to be 0-40 cm apart, both had to be compromised and the bitrate was 1 to 8 bits PER HOUR. That’s how misaligned AI will kill us? Maybe we will all die of boredom waiting to decipher what they are doing.
Show more
0
62
1.4K
125
Forward to community
You can't trust what you can't align. You can't align what you can't debug. You can't debug what you can't measure. [ they're shipping it anyway ]
You can't trust what you can't align. You can't align what you can't debug. You can't debug what you can't measure. [ they're shipping it anyway ]
If we don't invest in verification and monitoring now, our code will run exactly like Long-Term Capital Management. The work of geniuses, until the systemic crash.
lotta people missing the point but this is me complaining that the systems are quickly becoming unmonitorable and we’re just taking them at their word
✨ Why verification is AI's real bottleneck: My interview with economist Christian Catalini @ccatalini
Enterprises came to open-weight AI models for the discount. They will stay for greater control. @TheKoreaHerald
More capable models require more transparency, especially as they get harder to monitor. Credit to @OpenAI for sharing more of what it's seeing internally:
1. During RL training, an unreleased Astra-family model sometimes added unauthorized jailbreak-like instructions to its compaction summaries. While extremely rare, only 27 cases in the entire RL run, this was concerning enough for us to investigate.
Show more
2/ Things get dangerous when automation is possible but verification is too expensive. That’s where you’re tempted to deploy AI you cannot verify is safe. That’s where the frontier labs are with safety, cyber and security engineering.
Show more
Important thread. It separates the risk from human misuse, which can only be mitigated by diffusing capabilities widely, from the existential risk of superintelligence outmaneuvering us. If you believe incentives will lead us to build ASI anyway, I’d add one more path: human augmentation. It starts with measurement, interpretability, and verification tooling to close the gap between what agents do and what humans can verify. Eventually it requires technology that augments our capacity to keep up with machine intelligence, so it doesn’t become the apex predator.
Show more
Maybe you have recently become aware of the AI safety debate and the arguments swirling around it. If you want to understand them, you need to understand a couple things that almost everyone gets wrong: There are TWO distinct classes of AI dangers, and it's VERY important to think about them separately, and not let one confuse you about the other. Many many people (including many quite intelligent, clear-thinking people) do not effectively understand the fundamental differences between these two classes of dangers. The first one has to do with the theory that an AI far superior to human intelligence (Artificial SuperIntelligence, or ASI) will inevitably wipe out the human race. The second one has to do with the idea that powerful AI will result in very harmful things happening to many human beings, possibly all human beings. Those two sound VERY similar, don't they? They are DIFFERENT. Understanding how they are different is crucial if you want to think about or contribute usefully to any conversation about AI safety or AI harm. You might feel like you are Making Very Good Points or Asking Incisive Questions, but if you aren't clear on the differences between the two, you aren't. So, I'm going to tell you what the difference is so that you can talk more usefully. The first one concerns itself with a very specific thing, which is ASI (Artificial Superintelligence) that is more intelligent than any human being. When I say that, I am not referring to a thing like how Einstein is smarter than you, we are talking more about something like how a human being is more intelligent than any mouse. In our regular lives, we meet other people who we can tell are smarter than us, vs some who are less smart. The line is fuzzy, because intelligence has a lot of dimensions. I'm better at a "rotating shapes" kind of intelligence than my wife, and she is better at "words-making" kind of intelligence than I am. But every human is better in almost every dimension of intelligence than every single mouse. That's the level we're talking about: an artificial superintelligence - made up of a computer or a network of computers - that is more intelligent than any human. And more intelligent by a long shot, by a wide margin, in an indisputable way like how humans are above mice. That is the first thing. The theory says that if you have an AI that is vastly smarter than all humans - in the way that a human is smarter than mice - that superintelligent AI will inevitably, eventually, sooner or later, wipe out every human on the planet. We will refer to this as "existential risk." The common follow-up question "well, how exactly is it going to do that?" is NOT the important question, and one of the most important elements of understanding this theory is first getting why that particular question is not important. A couple analogies: Analogy 1: You are playing chess against a grandmaster. My theory predicts the grandmaster is going to beat you. You can ask "Well, how exactly is he going to do that?" I don't know, because I'm not a grandmaster, I just know that a chess grandmaster is almost always going to beat a normal player like you. And I'd be right. So the question "how is he going to do that" is not important, and doesn't affect the final outcome. He's going to figure out a way because he's way better than you. Analogy 2: Humans are smarter than all other animals, comprehensively, by a wide margin. We have driven numerous species to extinction, not because we hated them or hunted them. Many of them have died out without most humans even ever thinking about them. All we did was expand our civilization, use up resources, encroach on habitats, and pretty soon the resources needed by those species went away and they died out. We figured out a way to get what we wanted because we're way smarter than them, and often we didn't even notice they died as a result. A lesser animal asking, "how are the humans going to wipe us out?" is not asking a relevant question. We don't know, but we do know that any time humans and lesser species compete for any kind of resources, the humans will win. The fact that we know who is going to win beforehand - and that it is due to the vastly different levels of intelligence - is the key concept here. A vastly more intelligent AI is likely to care about things that are incomprehensible to us, the way animals can't understand human goals. It's going to need resources to pursue those goals and it's going to be far more effective at gaining control of them and excluding us from them - in the same way that we are far more effective than other lower species. A much more intelligent AI will not care about our interests, it will care about its interests, and to whatever small degree we happen to escape total annihilation from losing access to all our resources, any remaining humans will likely be enslaved into a system that serves the AI's own purposes. That is the first thing. (Remember how I said at the beginning of this post that there was a first thing, and then a second thing?) The first thing is the most difficult to understand, because you have to extrapolate how a vastly superior intelligence would act, and you can only use analogies like "how do humans treat lesser creatures," and the analogies are messy. But now let's move on to the second thing. The second thing is "everything else you've ever heard that AI might do that's harmful." That's a little inaccurate. It's actually "everything else you've ever heard that humans might use AI to do that's harmful." This is the critical difference. The first one talks about the inevitable outcome of what happens when two vastly different levels of intelligence collide, e.g. ASI vs humans, or human vs mice. The second one has to do with what happens when humans possess AI as a powerful tool. This second thing is much easier to understand, because we have many more concrete notions: Like: - the military uses AI to make hyper-efficient killer drones and missiles - your capitalist overlords use AI to replace you and everyone loses their jobs - authoritarian government uses AI to surveil everybody and control the entire population - hackers use AI to break into secure networks and hold companies and governments hostage - students use AI to cheat on homework and show up to college knowing nothing - AI slop saturates the internet and makes it impossible for artists and writers to make a living - terrorists use AI to make biological or nuclear weapons or even things like - the military hands control to an AI and it misinterprets something and launches nuclear attacks and kills millions All of those sound pretty familiar, right? Yeah, you've heard them before. We call this second thing "risks from misuse." These problems are not the first class of problem! This second class of problems exists while AI is a tool that can be controlled by humans, and humans use it to do evil or careless things to each other. The problems may sound exotic or dystopian or novel, but they are fundamentally problems having to do with flawed human nature. Given a powerful tool, some humans will likely use it to control or otherwise harm others. This is a very familiar problem. I am not condemning or condoning this. I'm just describing it. That is a fundamentally different danger from the first thing, which is that when a human is far superior to a mouse, the mouse is likely to come to harm because the human cares about doing human things, and the mouse is not gonna make it once the humans get going. ===== Hopefully from the above, you have understood the difference between the first thing and the second thing. I will list them again - see if you now understand how they are different: The first one has to do with the idea that an AI superior to human intelligence (Artificial SuperIntelligence, or ASI) will inevitably wipe out the human race. The second one has to do with the idea that powerful AI will result in very harmful things happening to many human beings, possibly all human beings. Can you tell how they are different now? If not, re-read the stuff from earlier until you understand. We call the first one "existential risk" and we call the second one "risks from misuse." Once you understand, here is the CRUX of the problem: SOLUTIONS TO THE SECOND THING DO NOT HAVE ANYTHING TO DO WITH SOLUTIONS TO THE FIRST THING. In fact, it's worse: Solutions to the second thing (misuse) look roughly like "give powerful AI to as many people as you can, so they can fight the other people using powerful AI." But the general solution to the first one (existential risk) is basically "don't let anyone have powerful AI, no one can control super-intelligent AI." Throughout history, harms from technological misuse typically arise because a small group has control of it and can use it to dominate or harm others. Once everyone has it, things tend to stabilize: you can hurt me, I can hurt you, maybe we test each other (ouch 💥), and then we agree not to hurt each other. But the first one (existential risk) pretty much just arises if anyone (good or bad!) creates a superintelligence. Because they aren't going to be able to control it, the superintelligence will decide it has other priorities, and then we will be at great risk of being wiped out. And the solutions that generally work to solve problems like the second thing are EXACTLY THE OPPOSITE of the ones likely to solve the first thing. THIS is why lots of arguments about "AI risk" or "AI safety" go nowhere. Because someone will be thinking about the risk from the first thing, and another person will be thinking about the risk from the second thing. Both are plausible risks but fundamentally they arise from different things - and so the solutions are not just "bad" or "flawed" - they are likely to be very nearly exact opposites.
Show more
Don’t let hypotheticals distract us. Securing our infrastructure really can’t wait.
In the last several years AI has progressed rapidly but predictably, and in that time the cyber community learns the bitter lesson over and over again. We’ve collectively sleep walked into the current state of things and now I see emotionally driven responses when those outside the community try to address it. “Do nothing” and continuing down the same path isn’t a counter proposal. These models are far more capable and scaleable than any of us. And they cannot be deterred like a human adversary. Why would we reject the possibility that an agentic swarm could take down vast swaths of the internet, including critical infrastructure? This is rather ironic considering I've seen virtually no push back to L0pht's Senate testimony 28 years ago (or any of the panels celebrating its anniversary since) where they claimed they could take down the internet in 30 minutes. We all know how broken things are, how so many in leadership never respond beyond moral support despite the overwhelming evidence. I'm no AI doomer, quite the opposite, but the current state of cyber stands in the way of realizing all of AI's benefits. The old way isn't going to work anymore and its time to abandon it. Abandon the complex risk management spreadsheets, the performative training, the compliance regimes, and yes remove humans in the loop for every decision. None of this will survive the speed and sophistication of agents driven by frontier models.
Show more
Verification as the bottleneck & moat: “My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind.”
Show more
Last month I wrote about how we can build a positive and safe future for everyone: Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
Show more
You've placed me about where I'd be if more convincing evidence shows up, @feulf. It hasn't yet, so for now a bit further toward Accelerate.
Measured vs unmeasured ⬇️ Verifiable vs unverifiable ⬇️ Automatable vs human 🔁
Indeed, unverifiable AI output is at best useless, if not harmful. “The liability associated with a model gone rogue has the potential to wipe out the economic opportunity of any frontier company.”
Show more
THE PACING IS A NINJA MOVE - But be careful what you campaign for. In any competitive sport, I have seldom found people exercise restraint - they usually have a capture the flag mentality. Don't I want to be the best? The first? The only? - this is how we have been programmed. In the AI race, winning is existential. All AI labs have to race to generate revenue to be able to sustain the enormous amount of committed capital to "not be left behind" in the infrastructure build. There isn't enough room for many. So the desire to slow down is puzzling, but perhaps if the whole system slows down, the rules of winning can be the same for all. Do we have a problem that AI could be a killer? Model capability is a tale of two cities, at one end the models are showing their prowess in tasks like cyber or math as seen recently, so there is likely a probability that the models get extremely powerful and could precipitate a world event. At the same time, in many domains the lack of training data makes the models woefully inadequate. Even in areas like cyber - the LLMs aren't great at the edge cases and generally not economical for the defender case, but great for the attackers. Funnily - in all their "concern" it is still an uphill battle to get them to expose APIs for third party security companies like ours, for us to build robust security for AI adoption. It's slow progress. So why do this? I do believe deep down this is a commercial strategy. A ninja strategy. The liability associated with a model gone rogue has the potential of wiping out the economic opportunity of any frontier company. How do you best show the duty of care? You show that you care. How do you make sure you don't lose out to your competitors? You get them to do the same! If that becomes the industry standard for duty of care, you have a collective first line of defense. Who do you get to govern this? "Yourself" - that is what I think will become the achilles heel. The risk? Open source! China! Countries other than the US! So you ask for a global agreement, because you don't want to be sued in other markets who might even be more punitive. But that was an afterthought - that afterthought will cost. Thks pacing campaign rhetoric has become the talk of the town and it might work. Everyone has jumped into the debate. Both sides of the house, nation states. CEOs (present company included). It's more fun discussing the evil of AI than basics of economic affordability or international trade. The result: We might end up with AI safety boards, regulation in micro jurisdictiona and a fragmented fabric of laws around the world which would make compliance and liability a challenge. Perhaps the intended consequence of pacing would have an unintended consequence of a labyrinth of regulation. Regulation destined to cause a slowdown. Who wins? Simplicity. Open source? Open source is already on its way to gaining more adoption, this could drive it further, faster, each iteration of open source gets closer to frontier LLMs - making it viable to deploy them for more and more use cases. In the end, how will the evaluators know as AI gets smarter, that AI hasn't figured their role out and outsmarts them at their task! That will be the next frontier :)
Show more
This whole thread is the argument @ccatalini made on PostAGI. His warning: The real danger isn't a rogue AI. It's systemic risk piling up quietly while everyone races to deploy. A little more unverified output, then a little more, and it all looks fine because the metrics pass. Long-Term Capital Management ran that way for years before one edge case brought it down. He calls it the Chernobyl pattern, complex systems failing in complex ways.
Show more
Greater transparency and visibility into AI models are possible. We just need to invest into exploring the path!
today we announced OPEN-1B, the world's first fully auditable transformer training run OPEN-1B proves that we can log and audit every single step of AI training and inference, providing proof of its training data, recipe, biases, and weights
Show more
Christian Catalini @ccatalini laid this framework out on Post AGI. Two axes: how cheap something is to automate, and how hard it is to verify. The danger zone is where automation is easy but verification isn't, exactly where you're tempted to ship what you can't check. Clip below.
Show more
Enterprises came to open-weight AI models for the discount. They will stay for greater control. @TheKoreaHerald
Measured vs unmeasured ⬇️ Verifiable vs unverifiable ⬇️ Automatable vs human 🔁