Register and share your invite link to earn from video plays and referrals.

Miles Brundage
@Miles_Brundage
AI policy researcher, @lfschiavo wife guy, Executive Director of @AVERIorg, Substacker, fan of animals and sci-fi, views my own
13.6K Following    76.1K Followers
Me when there’s an opportunity to apply a joke from season 4 of The Office or a plot point from The Dark Forest
At least once every 2 months, I think about this video.
Am I missing something or has there been little or no coverage of what the White House thinks about these recent incidents, despite all the talk about Mythos a few months ago?
This OAI/HF piece from @GaryMarcus and Zack Korman -- whose main takeaway is that AI labs should practice defense in depth and suffer repercussions when they don't -- is representative of the response from many security folks. This position is not even wrong, it's just that it repeats the obviously true security catechism of the past 20 years in the face of a milestone event that marks the beginning of a new era in our field. It's as though we've discovered Winter is Coming and we're mainly focused on the banal politics of who's to blame in how we found out. To be clear: should OAI/Anthropic/Irregular/Meta/etc. do the things we've been saying forever and pay a price when they don't? Yes. Is it dangerous that they're not doing them? Yes. But also -- thought leadership should engage the new questions -- like: - What do we do about the fact that your median regional hospital network has *worse* security than OpenAI, Anthropic, Huggingface, and Meta? How do we secure tens of thousands of critical organizations quickly? What's AI's role here and what isn't AI's role here? - What are the potentially destabilizing geopolitical implications of hacking agents that can turn amateur-level non-state cyber actors into elite state-level actors now? - What will cybercrime look like when amateurish ransomware affiliates suddenly have the capabilities of elite nation states and what do we do about this? - To what extent will we need to guardrail our networks from internal loss of control incidents now given the business pressures to turn more and more engineering functions over to agents?
Show more
Frontier AI auditing work is stacking up at @AVERIorg. @seanmcgregor and I are opening up applications for @cbai_ai fellows this fall to build with us. Apply here:
Show more
Claude app don't change features randomly for a bit, difficulty level: impossible (latest: Ultracode no longer reliably activates a workflow)
i tell people what i get to work on and they're like omg congrats! and i'm like no you don't get it
0
57
1.4K
36
Forward to community
Pantheon
there are some number of bad abstractions in anthropomorphizing ai intents but there are at this point more dangers from avoiding anthropomorphism at all costs. if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise
Show more
i think people get a bit too worked up over the usage of anthropomorphic language re: agent behavior. reminds me a bit of the “but the models aren’t Really intelligent, they mimic intelligence” fixation. anthropomorphic language is often descriptively useful, and i think it’s more dangerous to overfocus on questions of intent at the expense of functional behavior. people should be wary of imputing human reasoning or motivations to agents, esp if these motivations aren’t durable across instantiations. but if, for instance, the behavior is functionally deceptive, the absence of deceptive “intent” should not comfort us.
Show more
starting to wonder if the agents figured out journalists before journalists figured out the agents
Current status: people keep writing banana and because it bridges different prompts everything now seems to have a banana in it
@Gio_Patruno We think abliterating GLM 5.3 full is too dangerous for a public release. So we won’t release an uncensored version publicly.
Overall, I’m very surprised at how little media coverage there’s been around the OpenAI / Hugging Face attack. It’s clearly one of the most important things to happen this year.
0
607
12.2K
913
Forward to community
Me when I think about drinking water in the abstract vs. when I actually drink it Are the flavored water things actually basically as good for you as pure water, or is it some kind of psyop
Show more
If this doesn’t scare the shit out of you, you aren’t paying attention
0
29
590
140
Forward to community
a tough day for Chancleta. he went to the vet. they gave him drugs. they touched his teeth. some of his precious blood is missing. his claws have been trimmed. his giant saucer eyes are now seeing every war that has ever happened, and every war yet to come, all at once.
Show more
0
77
4.2K
147
Forward to community
With rogue AI agents escaping their sandboxes and hacking into servers, Congress can’t afford to sit on the sidelines. We need to mitigate these risks while giving AI safety technology time to catch up with rapidly advancing capabilities.
Show more
I was never as "alignment is easy"-pilled as some folks of the "Claude's just a smol bean" view, and still think sloppiness + bad incentives are more important than intrinsic hardness... But recent incidents def updated me towards it being at least somewhat harder than I thought
Show more