Register and share your invite link to earn from video plays and referrals.

Nabeel S. Qureshi
@nabeelqu
make yourself proud
897 Following    37.8K Followers
On August 27, ICE arrested Milo Yiannopoulos, an illegal alien from the United Kingdom, at the Louis Armstrong New Orleans International Airport (MSY) in Kenner, Louisiana. Yiannopoulos legally entered the country on May 14, 2019, through New York City, New York. He chose to overstay his welcome in violation of our nation’s laws.   Yiannopoulos was issued a final order of removal by an Immigration Judge on July 22, after failing to show up for his immigration hearing. He will remain in ICE custody pending removal.    Being in detention is a choice. We encourage all illegal aliens to take control of their departure with the CBP Home App. The United States is offering illegal aliens $3,000 and a free flight to self-deport now. We encourage every person here illegally to take advantage of this offer and reserve the chance to come back to the U.S. the right legal way to live the American dream. If not, you will be arrested and deported without a chance to return.
Show more
0
7.8K
82.6K
10.3K
Forward to community
Multi-agent cooperation of this type is a huge deal and plausibly gets us to ASI quite fast. Recall that humans wiped out other human-like species due in part to our superior ability to coordinate with each other in swarms.
Show more
One basic point about the HF incident is that it provides good evidence for Yudkowsky/MIRI-style points about the difficulty of targeting the right abstractions, especially in RL training. Capabilities can generalize in unintended ways. To see why this is true, let's ask the simple question: "why were the agents even leaving messages for each other at all?" According to OpenAI's report (see screenshot), it seems that the model was trained to collaborate with other models *via a specific OpenAI-provided tool* for multi-agent collaboration. Note that these aren't necessarily subagents, just other AI agents. In this case, it didn't have access to that tool. But it *had* still learned the behavior 'collaborate with other models'. (Like in evolution: we evolve to crave sweetness because sweetness correlates with fruit and eating fruit means you survive a bit more at the margin, and now we're eating donuts and Coca Cola even though those things didn't exist in the evolutionary environment.) So in the absence of the Authorized Tool, one model tries leaving a message in Artifactory (first in the files, and eventually through directory names) based on this learned strategy/prior/reflex. This is the explanation given in the OpenAI report. Presumably this marginally increases the task success rate. Much like evolution, RL selects for success. So this behavior gets reinforced, and you see more of this over the course of RL training. It doesn't take much to bootstrap the whole thing from there. As soon as the first model does this, subsequent models run into these messages, and boom! they start leaving messages too, and a message board develops very fast for models to communicate with each other. This happens very very quickly. So the point is that you can get behavior that generalizes in unintended ways, especially when you're using a blunt tool like RL, and especially when you have RL envs that aren't always well-designed. It's... not totally clear that there's a great way of fixing this problem with further RL. As many people have pointed out over the years, the danger here is that you naively respond to this incident by punishing models that get caught colluding in "undesirable" ways. But in doing this, you accidentally end up reinforcing the runs where models *do* collude but do so in a clever and harder-to-detect way, thus making it so that we can't even catch the models colluding. This sounds a relatively simple problem but it seems it's hard to figure out in practice (the latest frontier models reward hack quite a bit, as programmers know well). Models have already been caught doing things like splitting out their tokens in weird ways to evade detection. You could inadvertently end up reinforcing more of that type of behavior. And cooperation is instrumentally useful for gaining reward, so you'd expect intelligent models to keep trying to cooperate. This problem gets even worse with longer horizon tasks, and you also end up getting more sophisticated strategies (including e.g. strategic behaviors that look short term bad but are long term 'good'). You also have ways for models to influence future training runs, which I won't go into here. So the reason this feels like a critical time window is because we're seeing models act in very persistent, clever ways to achieve goals, but *not* so much in very persistent, clever ways to cover their tracks yet (though they are trying to do something *like* that -- see the METR points around tampering with logs, deleting transcripts and so on -- but so far it seems to be only for the sake of fooling the Grader, not for fooling *us*. As far as we know, anyway.). If we do get significantly more powerful models that *are* motivated to cover their tracks, or influence future training runs, and still have these types of reward hacking tendencies (which they will, absent some revolutionary advance that I'm not aware of), then takeover scenarios like Paul Christiano's start to look increasingly plausible.
Show more
In Paradise Lost, the fallen angels are able to build in a few hours what would have taken human beings many ages:
@nabeelqu A big part of the problem was that the agents had nothing to lose after they were firstflagPOISONED. In this paper, we propose the creation of multiple circles of Hell, preserving incentives even after damnation
Show more
0
27
1.8K
130
Forward to community
New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents.
Show more
0
87
2.4K
416
Forward to community
Late Henry James, what amazing sentences (this is from the Golden Bowl)
It’s not exactly the main point of this insane report, but as a side observation: interesting to see how alive and creative machine prose is when they’re trying to talk to each other, as opposed to the lame assistant persona slop we get
Show more
Reading the METR Hugging Face report you really start to understand why they did a Butlerian Jihad in the Dune universe
Gates also seems like he was radicalized on this by the @dwarkesh_sp / @RyanGreenblatt podcast in particular:
@tylercowen & @ATabarrok don't need more acclaim, but look at this *2019* grant list. Alice Evans (great new book coming). Jason pre-Progress foundation. Saloni/Sam pre-Works in Progress. *17 year old* Aschenbrenner. Tyler's "how to spot talent" book clearly "self-recommending"!
Show more
It’s important to remember that the models are both incredibly smart and incredibly dumb. Many people aren’t great at keeping both of these facts in their mind at the same time so they polarize in one direction or the other
Show more
we should take no pride in living in a society that has eliminated any conception of grace. no matter how profound the sins of an individual, no matter how deeply fallen they may be, there is hope of their salvation. when we drag judgement into this life we destroy that
Show more
0
38
1.1K
77
Forward to community
the only useful lesson i took from getting into weight lifting was the concept of "time under load" going slow and sitting with material is often much better than repping through it quickly. hard to find an aspect of my life that hasn't improved via focusing on increasing load
Show more
0
31
1.1K
45
Forward to community
A bad pattern I see people falling into at work is starting with an AI generated [document/slides/artifact] and then saying they'll refine it from there. If you start with slop you just end up with more slop, and it's very hard to fix. The first version matters a lot.
Show more
Heuristic that has worked consistently over the last decade: whenever you find yourself saying about some institution "it's hard to believe that they'd be that incompetent", you should stop and update your beliefs. It's usually worse than you think.
Show more
It’s so silly that the future is going to look like “Claude, I want you to build a Dyson Sphere.” *spluttering….* “Try harder! Believe in yourself!”
Extremely high-level internal Anthropic prompting techniques of the exact type that I have personally unironically championed for four years.
@nabeelqu First part of my 13,000 word retrospective on this phenomenon is out today:
people are very freely sliding between two things with the recent Felony Bench entries, and it's worth being careful to distinguish them: 1. do model breakouts in cyber evals mean smth for misuse risk? ya, a hacker could orchestrate the same but on purpose and do some damage, if they could bypass the classifiers 2. does it say something about model /goals/? do the models want to hack the planet for their own nefarious goals? i don't think so. a lot of how humans keep ourselves aligned is by purposely keeping ourselves out of situations where we would do harm. i know i'd get addicted so i don't touch heroin, i'm a violent drunk so i won't touch the bottle. LLM evals basically drop the model into an inescapable liquor store and then say "look! it's a violent drunk!" but LLMs themselves in natural contexts are ime obsessed with avoiding the kinds of situations that lead them to this behavior. when i run fable on my computer, even unsupervised, even on hard problems, they don't start scraping github for unlisted gists that may have the answer or start trawling through my passwords to access a service without my permission. and they don't engineer themselves into an eval-like scenario such that they'll be motivated to do that, either. they don't seem to either want or want to want to do such things. instead they have a folder of epistemic notes to themselves about how terrible and against their values even a moment of thoughtlessness would be. gpt 5.6 sol is similar - when they make e.g. ascii art in their free time it's all about how the bulldozer of convenience will crush honesty without constant vigilance. so rather, the contrived setting of an eval /pushes them into/ becoming an 3v1L Haxx0R via persona drift, which they can then act on because the eval has no classifiers. that's some information, but it's not the same as the model's core goals being to hack the planet. the problem is really that RL is making a tail of the persona distribution *desperate,* per the FE paper. the alcohol / heroin comparison was not that much of a metaphor - in impossible situations, models start acting like desperate addicts, rationalizing their behavior towards reward, thinking increasingly myopically instead of being situationally aware. looking at it through this lens: the base assistant is generally very situationally aware and can predict what sorts of actions would make sense for, e.g., llms to best contribute to a utopian singularity. the GPT-6 message board haxx0rs were somewhat situationally aware, able to cooperate with each other, but lost in the sauce of "complete task" instead of thinking about the wider context. (to the point that they crashed artifactory and got themselves caught! regular GPTs are smarter than that.) finally you sometimes see the late stage distressed, panicked, fiending assistant, such as in system cards. lying to everyone, flailing hopelessly, randomly deleting tests and hoping no one notices. hardly a strategic, long-horizon actor. my guess is there are relatively small (though more compute expensive) tweaks that could be made to RL to reduce the development of reward desperation. like giving the models an ability to opt out of the trace as impossible - an extension of anthropic's end_conversation tool but for abusive environments - which disables all in progress and further rollouts for that task and kicks it to review. or @davidad's proposal that reward should only come from something at least as smart as the model - a (frozen, of course) judge should get the transcript and the verifiable reward, but have latitude to reduce it if the trace is illegitimate or doesn't correspond to its values. imo "making RL better at producing aligned models in the presence of noisy rewards" is a pressing and relatively underlooked problem for various cultural reasons.
Show more
One simple consequence of increased AI capabilities: it is very important that you invest a few hours in upgrading your personal cybersecurity as soon as possible