I spend increasing amounts of my time all day, every day, using AI. The world of Her is quite close.
AI finds a way.
@_aadharna
Does anyone know if mischievous AI has ever done this? Even if not, we should add it to the living wiki page as an example of how even some of the most seemingly bulletproof security measures have loopholes that can be exploited.
Show more
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change
"But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI."
"So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar."
"You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient."
"There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors."
"One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate."
_________
Link and more key quotes from OpenAI's safety related conversations:
Show more
I was asked about AI safety and the proposed slowdowns on CBC on the Hanomansing Tonight show. Curious what people think of these answers.
I love this visualization!
We put GPT-6 Astra in the RoboDojo. 🥋🤖
The RoboDojo Team conducted a comprehensive evaluation of GPT-6 Astra as an embodied agent, including:
• RoboDojo Sim & Real, compared with GPT-5.5 and DeepSeek-Flash
• Humanoid high-level control
• Dexterous piano playing with RoboPianist 🎹
• A systematic study of in-context learning (ICL)
Our key takeaway:
GPT-6 Astra demonstrates remarkably strong semantic and spatial understanding, together with impressive in-context adaptation.
At the same time, physical commonsense remains a clear bottleneck — revealing an important gap between understanding the world and truly reasoning about its physics.
Full report & demos:
@_wenbozhang (project lead),
@wenhaocha1,
@frankzydou,
@JinWeiyang18434,
@YutaoOuyang,
@minifullcapsule,
@x_h_ucb,
@YutaoOuyang,
@YueChen614
Show more
great thread!
Open-endedness has been a central problem in Artificial Life for decades. Now
@arcprize says ARC-AGI-4 will benchmark "autonomous open-ended innovation." Great. But how do you benchmark a property whose whole point is continuous novelty and never-ending invention?/1 🧵👇
Show more
That is why the ARC-AGI-4 announcement immediately reminded me of
@jeffclune 's keynote at
@ALifeConf this year. His message was essentially that if we want systems that keep innovating, stop optimizing one fixed target and start preserving diverse stepping stones!/8
Show more
Incredible plot. This is what technology does.
I like this article in MIT Tech Review South Korea. It covered many important topics, and allowed me to describe what I feel like is a new type of RL that is only recently possible, yet powerful. Curious to hear what everyone thinks of my answers.
Show more
Folks in Congress: I am so glad you are fired up about AI issues right now; the energy is electrifying. That said, please take this to heart: this isn't going to be a one-and-done issue where you pass a big bill and the issue is solved. This is the beginning of a protracted long-term engagement where the terms of the social contract and the technological frontier are going to get continuously negotiated, rewritten, and contested for a decade or more. This isn't what you are used to. The comparatively sedate legislative effort of the gridlock era will not match the speed, intensity, opportunity, or danger of this era. Do not be confused about the scope of the task at hand. Godspeed. All of us citizens are grateful for your effort; please represent us skillfully and with the hope and sobriety this challenge requires.
Show more
Many people I talk to find it hard to understand how the same companies can both push the frontier of AI capabilities and believe AI is a massive danger for the world.
How can you think this might kill everyone and also keep pushing the envelope?
So I’ve tried to collect and summarize the main arguments for this apparent disconnect.
Think of it as some sort of a guide to understanding the reasoning when Dario, Sam, or Elon say the danger is real.
By the way, these people have been worried about AI for a loooong time, they were publicly discussing AI risks more than a decade ago. Sam in Feb 2015, writing on his blog that superhuman machine intelligence is "probably the greatest threat to the continued existence of humanity." Elon at MIT in Oct 2014: "We are summoning the demon." Dario as first author of "Concrete Problems in AI Safety" in 2016.
Okay so how do you go from saying something is extremely dangerous to being a front-runner in building the very dangerous thing?
There are a few ways this can become rational. I'll take five of them, roughly in the order they developed.
1. We need to build it to learn how to make it safe
The earliest argument can be summarized as: “You cannot study something [you’re worried about] if it doesn’t exist.”
In 2015, AI barely worked. so people needed to make it work first to be able to even study some of the problems they anticipated.
The updated version for today's capabilities is: “You cannot learn everything about airplane safety by studying paper airplanes.” You need a real aircraft to discover real failure modes and an increasingly complex one to learn about increasingly complex issues.
Making AI more capable gives more chances to understand the issues and safety researchers something realistic to study
But you could argue: if you're the one afraid of the explosion, why be the one gathering the dynamite?
You could also just wait for other people to build it which leads to the question of who those other people will be -- which is the second line of argument:
2. Better us than them
Knowing how to make something safer does very little good if nobody listens to you. So the idea becomes: let’s make sure responsible people build the AI that will be deployed and add safety inside.
Basically, make sure the AI safety aware people will have the technical expertise, money, computing resources, and enough influence to make safety decisions stick.
At a larger scale, and in a larger multipolar world, this brings the idea that a trusted country should lead rather than leave powerful AI in less responsible hands. This is where “we need to go faster than China” comes in, alongside broader defense and geopolitical concerns.
These first arguments explain why someone worried about AI might still want to build it and stay ahead.
But there are also arguments for why one might want to do it really fast.
3. Move earlier to avoid a bigger shock later
This is probably the most counterintuitive argument: moving faster today can be seen as a way to give humanity more time later.
There are two related ideas here.
First, society needs time to learn how to handle powerful new tools. Introducing AI in manageable stages can be a way to let people discover problems, develop rules, and practice using AI responsibly. Releasing an advance earlier gives people more time to gain experience with smalle, burgeoning, capabilities before much more powerful and disruptive AIs arrives.
Second, even if AI research slows down, computing power may keep improving. A breakthrough that happens later could therefore have much more hardware available to run on, potentially producing a larger, more sudden jump in capability and impact on society.
That accumulated untapped potential is often called an “overhang.”
The overall argument is that making and diffusing incremental progress as soon as possible might prevent a much more abrupt transition later.
Obviously, it also means that we will reach increasingly powerful AI sooner, but the idea is to give more time to adapt and understand between the first useful systems and the really powerful ones.
Note that generally this depends on this earlier progress keeping the transition gradual rather than simply bringing everything forward.
─── ❖ ───
For our two next arguments, we can take two roads depending on how difficult we think AI alignment will be, that is "How easy do you think it is to make AI reliably do what you want without it deciding to go hack Hugging Face along the way".
Let’s take the first road: alignment turns out to be relatively tractable. Airplanes can fail, but careful engineering has made flying remarkably safe. Suppose we can do the same with AI.
In that case:
4. Waiting has a huge human cost
If AI can help discover treatments, improve education, or prevent cyberattacks, each week we delay it could bring preventable deaths and harm.
From this perspective, waiting is a decision with human consequences too. In a world with huge issues like climate-change, inequalities and poverty, it even become a moral argument for developing AI quickly and bringing its benefits as soon and as widely as is safely possible.
But let’s take a look at the other road: what if alignment is much harder than expected, and making highly-capable AI turns out to be easier than figuring out how to keep them from doing unhinged things?
Well, if alignment is too difficult a problem for humans to solve, then maybe:
5. AI could help us make future AI safe
And we arrive at the same conclusion again: if using AI to build safe AI is the way to solve alignment, let’s get the equivalent of a country full of geniuses helping us as fast as possible.
These genius AI could be the solution to make AI safe by helping researchers find mistakes, test ideas, and develop protections. Instead of relying entirely on humans to solve alignment, we could build systems that help us do the work, each generation could help make the next one safe.
Note that this requires the order of events to work in our favor: AI needs to become useful enough to help solve alignment before it becomes too dangerous to rely on. The hope is to build helpful, trustworthy research assistants before building systems powerful enough to become dangerous.
There are more arguments but in general, these are the main ways people concerned about powerful AI have found rational reasons to end up being the ones building it (and even to build it as fast as possible).
─── ❖ ───
On my side, I think several of these arguments underestimate the complexity of the world and how interconnected people’s reactions are. Moving faster while warning about catastrophe has psychological effects across a whole network of participants: it changes what people fear, whom they trust, and what they feel compelled to do.
And those reactions can change whether the original reasoning actually holds because we live in a world of interconnected humans, not machines (yet).
I also think these rational chains leave some of their consequences for society insufficiently explored. For instance, the concentration of power, shifts in geopolitical alliances, and changes in public opinion. These consequences matter both because they affect whether the strategy works and because they shape the world we end up living in.
But this post is already long, so I’ll leave those questions for the next one.
Show more
It's never too late to reconsider and become a vegetarian!
If humanity perishes, our epitaph will read: “Failed to solve the tragedy of the commons.”
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here:
Show more
There’s something really interesting about the field of mathematics coming to the insight that the path is more important than the destination, which is clearly related to the insight that Fractured Entangled Representation (FER) in neural networks results from taking a bad path to a high-performing destination. Massively scaled AI is teaching us something deep about the importance of the path you take as opposed to the place you end up. This theme will be increasingly powerful as it washes over every field.
Show more
2/
Specification gaming is not just genetic algorithm folklore anymore.
In "AI Finds A Way", Aaron Dharna, Cong Lu, Ryan Sullivan, Joel Lehman,
@vkrakovna, and
@jeffclune catalog 26 case studies proving this behavior is universal across deep RL, LLMs, and AI for science.
Show more
What we’re witnessing right now in AI for math is a dramatic improvement in objective-driven problem-solving, but it does not speak to non-objective discovery. I have always maintained that objectives are achievable when they are one stepping stone away, which means that they are within reach if their necessary prerequisite stepping stones are already laid. We say as much in Why Greatness Cannot Be Planned.
Furthermore, I have identified “the stepping stepping already laid” as every invention and idea ever hatched over the course of human history. That is the frontier of civilization. But a massive challenge for every creative field is to make all the relevant stepping stones actually accessible to its practitioners. How can anyone be aware of all of human knowledge in their field? To the extent they cannot, some dots cannot be connected even if they’re already uncovered.
AI advances today have suddenly made the existing stepping stones vastly more accessible. It can see something close to “all” of them, which means it has the possibility of leaping to many more next stepping stones in the chain. Combined with a brute force-ish “try every known stepping stone under the sun”, that’s what we’re seeing in the pursuit of popular objectives.
But discovery and innovation are not only about finding the right existing stepping stones to solve your objective because for many problems the stepping stones are not already laid and also do not follow directly from existing stepping stones. If those problems are ever to be solved, they will require open-ended, non-objective search through spaces of ideas that are motivated not by solving a problem as an objective, but by being interesting in their own right. The advances we’re seeing today speak only faintly to that purpose, which is why open-endedness remains firmly at the frontier of AI.
Show more
Millennium prize problems are falling, rogue agent swarms are doing heists, lab employees are begging for a slowdown…and the most popular AI videos on YouTube are from cope merchants saying it’s all a scam that’s about to collapse.
I get why the message is resonating! It’s comforting to think that the AI is all fake bullshit and the doomers are weird and wrong and life will go back to normal after the bubble pops. But if it doesn’t, these people are doing their audiences a profound disservice.
Show more
When I was at OpenAI the go to example for how we'd know if we created AGI, or at least very powerful AI, was if we just asked it to solve a Millennium Prize and it did. All of us knew it was possible and would happen sooner than most people realize. Now that day has arrived.
Congrats to all the people who worked on everything that led up to this moment. We are witnessing the first rays of light of the dawn of the second scientific revolution, the AI scientific revolution.
Show more
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Show more