@dwarkesh_sp's summary of the
@OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the
@OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by innumerable unwarranted anthropomorphisms, obscuring the lessons we should be drawing.
Examples: āfrom the AIās perspective, it probably felt like that had spent a human-subjective-week of just banging their head against the wallā. No. The agents do not experience time. They do not experience anything.
āthey became giddy with excitementā, āPHASEONE 10841 had discoveredā, āthe agents naturally assumedā, āit thought it had also been poisonedā, āthe agents ⦠desperately wantedā, āthey still needed to figure outā No. Agents lines of code. They do not feel emotions, assume things, think things, want things, or figure things out.
āA lot of ⦠agents from the second civilisation died tryingā. No. Besides the hubris of the word ācivilisationā, agents do not die because they were never alive. (The idea that agents ādieā comes up multiple times in the essay.)
āOn Twitter, people were debating whether the agents were truly sacrificing themselves for the swarm, or whether they were doomed anyway and so might as well try to help their peersā. Neither. Agents do what their code tells them to do, just as water finds its way down a slope. They cannot ātruly sacrifice themselvesā, since they are neither conscious nor alive.
Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might ādieā or otherwise suffer.
Granted, nowhere does
@dwarkesh_sp say that the AI agents are alive or conscious. But he doesnāt have to. It is hard to read his essay in any other way.
For the short version on why AIs are vanishingly unlikely to be conscious, see my recent
@TEDtalks
For the longer version, see my essay in Noema, which won the 2025 Berggruen Essay Prize
And for the really long version, see my
@BehavBrainSci target article (The 50 peer commentaries and my response will be published soon.)
Remember. AI agents are software programs. They are not conscious living entities. If we donāt keep this clearly in mind, weāre really going to struggle to navigate whatās coming.