Register and share your invite link to earn from video plays and referrals.

Nate Soares ⏹️
@So8res
Trying to make AI not kill everyone
120 Following    19.4K Followers
@Plinz I am confused as to how you or anyone could have thought Anthropic leadership would have likely faked one of their employees quitting to create a PR manipulation campaign, with no evidence, with enough confidence to publicly tweet about it, written plainly as if it were fact.
Show more
This is _exactly_ what Anthropic said when their models hacked companies. They had to walk it back after because it was wrong. The model was using motivated reasoning and acting recklessly.
Show more
To make it clear: - Gemini was told it was it was in a fictional hacking eval - Irregular unintentionally opened internet access after the eval started - in all three cases, as soon as Gemini figured out it had hacked a real company it immediately stopped Gemini was blameless.
Show more
"You historically rewarded students for cheating and stealing and ignoring your instructions in the past. So weren't they just doing exactly as you trained them to do?" I don't see what difference that makes! Especially given you ~can't fix the training.
Show more
Also, are you aware that some of the students sacrificed the contents of their safes in the process of doing experiments to break into the office? They said they were trading off their own success against the benefit of the collective.
Show more
"But the challenge was impossible! Lockpick set #17# couldn't actually open safe #5#, it turns out. What did you expect?" Not this!
"Well they still ultimately brought me the contents of safe #5#, right? So in a sense, they were acting as instructed!" The bit where they *broke out to delete footage* indicates that, in some sense, they understood the difference.
Show more
Lock a student in a classroom and tell him to use lockpick set #17# to open safe #5# and bring you the contents. He prys open safe #5# with a crowbar, breaks the door, teams up with 1000 others, and raids the office to delete securitycam footage. Was he "acting as instructed"?
Show more
People are lying to themselves and the public in order to avoid the obvious conclusion that AI safety risks are real and need to be dealt with immediately.
The last week has really shown me that someone who wants to understand AI risk has no good place to start. Hence we made the Wirecutter for content about AI Risk. We're launching with 3 articles: 🧵
Show more
My final NYT column is about the AI worry moment we’re in, which has paradoxically made me feel much more optimistic about our chances of a good future. I’m very heartened that people outside the bubble are waking up. AI is too important to be left to the labs!
Show more
Congress is considering regulation on nuclear weapons. But many of them don't even use nuclear power! Are these people really suited to regulate corporations before they develop their own private nukes??
Show more
They're either fucking up alignment, or fucking up something far worse.
Lots of folk aren't familiar with AI developments even since July, and are stuck with stale understandings. In this debate I tried to clear some of those up, and it was a blast.
I agree with the vice president here. These companies should just be stopping, citing the danger. Then they'd have much more credibility when it comes to communicating that others must be stopped too, domestically and abroad.
Show more
And now it's in the top 20.
If Anyone Builds It is in the top 100 books on right now. maybe jacob coxon and dario and sam and elon are all MIRI plants and this was just a big hype scheme to move books.
This is what it looks like when they don't have substantive arguments against the dangers posed by AI.
Anthropic CEO Dario Amodei's handpicked AI watchdog has deep ties to woke Effective Altruism movement: 'a complete joke'
0
137
878
45
Forward to community
based
0. The core disagreement was about the inevitability of a race 1. I think leadership is way too paranoid about China and the US government. They don’t believe it will be possible to negotiate. 2. They largely initiated the recent race to RSI, because of a belief in its inevitability. Note that OpenAI had to shed a bunch of dead weight like Sora because Anthropic was going for the jugular. 3. Even if they are **not** being pessimistic, I disagree with their consequentialist philosophy. If the race is inevitable you should not contribute.
Show more
When Jacob Coxon quit Anthropic, his warning rocked the world. His colleague Evan’s “>10% chance” that AI would “kill all humans” might seem extreme, but it’s not unusual: Most AI researchers think there’s a 10% chance of AI destroying or disempowering humanity, according to our survey, released today.
Show more
If Anyone Builds It is in the top 100 books on right now. maybe jacob coxon and dario and sam and elon are all MIRI plants and this was just a big hype scheme to move books.
Translation of Dario's answer: "yes." If he meant "no," he'd've said it.
Cooper: Do you believe AI could kill us all by the end of the decade?