I agree with Dario. If we are to survive, we must pace the frontier. The speed of development is quickly becoming too fast for us to keep safe.
i'm late to the party but: from his time at OpenAI I remember Jacob as a very thoughtful researcher and he continues to be so in this thread. neither anthropic nor openai are on track to solve alignment to a degree sufficient for shipping superintelligence and we need to slow down
Show more
I think we have a ~50% chance of all dying as a result of superintelligence, mostly due to actions in the next few to 10 years. I don't expect to have that resolved to below 10% or above 90% before we either make it through, or we don't.
We will have to act despite uncertainty.
Show more
I cofounded one of the largest compute companies in America.
I want to publicly state that I support regulation on the development of AI - even if it means slower progress for my own company.
1/6
I worked at Google DeepMind and now at Anthropic. This is a common sentiment amongst my peers.
(I write this in personal capacity.)
There is not yet a viable scientific plan to solve risks from recursively self-improving AI. Please look up!
Show more
Jacob was a senior researcher who joined Anthropic in May. My Anthropic colleagues and I had been trying to recruit him for ~2 years, because we knew he was a strong researcher at OpenAI. Before he left, I pitched him to stay and join my team, and I was sad he decided to leave, as are many of my colleagues. 100% agree with him that AI poses serious risks to society, and I'm glad he's speaking out!
Show more
Former Anthropic employee Jacob Coxon says the AI industry is "gambling with our lives." He tells Anderson what he finds "most scary is if AI is used to make itself more intelligent" and warns these companies are "compelled to race toward building a deadly technology."
Show more
"We know how to control nuclear weapons... We don't yet know how to control AI."
Former Anthropic researcher Jacob Coxon, who recently left the company, is sounding the alarm on artificial intelligence, calling it "possibly the most dangerous technology that humanity has ever created."
Coxon, who says he spent three years conducting research at Anthropic and OpenAI, wrote in a viral post on X that the companies are more focused on beating each other to build the most advanced AI models than they are on safety.
He says an international AI arms race would be "disastrous," arguing global cooperation is the only way to manage the rapidly advancing technology.
@SpecialReport
Show more
“You just said that 10% doesn’t seem an unreasonable estimate that AI could kill all humans”
“Yes”
“Wow… oh my God.”
@vicderbyshire speaks to Nobel Prize Winner and so-called ‘Godfather of AI’ Geoffrey Hinton about predictions by an Anthropic researcher that there’s a more than 10% chance AI could kill all humans within a decade.
#
Newsnight#
Show more
I've known Jacob since he was a resident at OpenAI. He's a real person with real intentions to make the future go well, which landed him at Anthropic for three months. He quit for those same intentions, costing him a huge sum of money.
Show more
I'm seeing lots of dismissal of Evan and Jacob as being insincere or having ulterior motives.
I read about the extinction risks of AI in 2017 and was so alarmed I pivoted my career and moved to the US to work on AI safety.
I worked at OpenAI for 3.5 years. I know many of these researchers IRL. They have been talking about these risks for years. These are sincerely held beliefs.
You might disagree with the arguments, the probabilities, or think the benefits are worth the risks. But please don't tell yourself comforting lies that this is just a marketing ploy.
Show more
@jillgun I would burn my equity to the ground in a heartbeat for a 1% higher chance we make it out of this situation alive. I expect a great many of my colleagues across the industry would as well.
I promise you, we are actually just fucking scared, it's not galaxy brained marketing.
Show more
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet.
METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation.
Show more
It's time to sound the alarm - louder - on reining in Artificial Intelligence.
It's becoming more clear the threat AI poses to humanity, so I'm calling for immediate action from the industry and Washington.
Show more
[Writing this in a personal capacity, not on behalf of my employer (Anthropic).]
Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:
1. AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.
2. Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely.
3. Unlike traditional software, we can’t “program” AIs to behave how we’d like. AIs frequently severely misbehave. For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this.
4. We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them. Insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.
5. Many AI developer staff desperately want to slow down to figure out how to build AI more safely. That was the intent of this open letter (which I signed):
I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.
Show more
Reminder that we've proposed concrete options for pacing the US frontier:
And a more in-depth proposal for international coordination:
I don't know what my probabilities are on literal extinction, but I think there are a number of ways AI could go poorly for humanity, and at the current frankly terrifying pace humanity will be quite lucky if we manage to find and stay on the narrow path between all the bad outcomes.
I am heartened by the many costly actions OpenAI has taken recently (detailed in several recent posts), but regardless of what you think of OpenAI, this is not a problem that can be solved by any one company (or country) in isolation. We need coordination to be able to approach future capability increases with an appropriate degree of caution and humility, and we need it yesterday.
Show more
I left Google DeepMind in June. Jacob is right: many researchers believe they are building something that could kill everyone on the planet. It was literally my day job to think about how to stop that.
Show more
I’m very glad
@EvanHub is trying to figure out how these models work and sharing his results. And also, holy shit, we are in a scary timeline.