Until AI companies have adopted actual preventative measures, I'm afraid we are prone to many more Hugging Face incidents.
Two big things that still seem missing:
1) Do not allow AI to make **any** code-merges into your control system unless they've been affirmatively cleared as safe.
- If you don't pre-clear these code changes, they can *turn off* your other controls. This can make any of your after-the-fact responses, like paging, totally ineffective.
- You should be using a high-quality monitor with CoT access for this. It is not enough even if you have a human review every pull request; you want to check the transcripts for misalignment evidence, and they're way too long.
2) Make sure your agent automatically fails-closed in very serious cases. i.e., implement a circuit breaker for if your monitor is getting blitzed by a bunch of attacks.
- Otherwise, the attacker gets a very high number of retries, and can attempt very harmful things, so long as it is confident in sneaking _one_ through before humans can respond.
We describe more inside Principle Four of Guidelight's Control standard: "Stop the AI from taking harmful actions even if it tried", but these are IMO the two most important practices.
I will have a hard time believing the industry has taken the recent incidents seriously, until these measures are in place.
Show more
They were just trying to accurately measure the model’s abilities, which is useful for _later_ interventions, like deciding what mitigations to apply (classifiers, additional refusals training, etc)
See eg:
This was a capability evaluation, not a propensity evaluation
Show more
Encode is awesome; highly recommend this role for folks in the thick of the AI community, who want to make sure the technology actual goes well for the world
Writing a sci-fi novel about a human thief who steals the take from his robot co-conspirators after a successful heist
I’m titling it “The Body Keeps the Score”
Ex-OpenAI safety researcher
@sjgadler says AI labs are monitoring their models like a bank that lets the robber turn off the security cameras:
"Imagine you ran a bank, and you want the bank to not get robbed. The way you do this is not setting up a security camera and every hour you check the feed to see if the bank was robbed in that time, and then you try to respond."
"You certainly don't leave it where a criminal could walk into the bank and turn off the camera, and that's the equivalent of what's happening at AI labs today, not just in terms of automated AI R&D and recursive self-improvement, but broadly across the board."
Show more
It's awesome to see OpenAI speaking out for laws that require preventative control measures.
Today's industry practices are much weaker than that bar, though. We score the best company just 3-out-of-5 on their preventative controls.
Show more
OpenAI Global Affairs put out a new statement today making clear that they don't think current state laws like SB 53 in CA and SB 315 in IL are sufficient and that its important for legislatures to revisit legislation in light of lessons from real world events like the HF incident.
I still don't love their obsession with constantly repeating the idea of "reverse federalism" (seems kinda like normal federalism to me), and their line about supporting SB 53 is a little funny to me given my recollection of their engagement on SB 53 (though perhaps they are saying that they now support it even if they were not fully supportive at the time), but overall this is a good clarification.
Actions will speak louder than words as always, but some of this language is pretty specific and I particularly appreciate the language that laws should be "requiring monitoring of frontier models under training or evaluation" - this should be a painfully obvious takeaway after some of the recent incidents, but it is still a somewhat novel idea to policymakers that models can pose severe risks even during training and its good to have OpenAI say that explicitly.
Show more
NEW from me in the Guardian --
I agree with the AI company employees who wrote, in their "Pacing the Frontier" letter, that the government should be stepping in.
But companies could also be doing more to step up on their own already.
I discuss 4 ways to prepare for a possible slowdown:
1. Auditing of each company is a key part of enforcing a slowdown. Companies can do pilots of that process today.
2. Actively participate in + set up new cross-industry governance bodies, e.g. to share safety lessons.
3. Invest in the technologies needed to verify a slowdown (or another kind of agreement) with China.
4. Push for legislation that starts putting key institutions in place.
I've been worried about race dynamics in AI for a long time, and e.g. coauthored one of the early papers on it. It was hard to keep up on safety and security in the GPT-3 days, let alone today.
The competitive dynamics are part of why I founded
@AVERIorg, and am trying to make frontier AI auditing effective and universal as quickly as possible.
But it's important to distinguish a difficult situation from an impossible one, and companies have a lot of agency that they aren't always using.
Read the piece here:
Show more
I'll be on MTS at 1 PT to chat about where AI companies are dropping the ball on controlling their systems, and Guidelight's work to change this.
See you then!
DATACENTER DISCOURSE | 0x ALPHA | DEEPSEEK FLASH VISION
Had a great conversation with Sarah about this moment in AI, Guidelight's Control scorecard, and the shape of the problem ahead. Check it out!
the Consistently Candid podcast is back!
in this episode, I chatted with
@sjgadler. Steven formerly worked on safety and policy at OpenAI and is the co-founder of Guidelight AI Standards, one of the organisations behind the recent Pacing the Frontier open letter.
we get into the open letter, Guidelight's new scorecard for AI safety practices at frontier AI companies, the recent Hugging Face hack, & much more!
link below :)
Show more
The Black Hat talk was the most must watch item of AI content this year, but imo this is #
2#.
Owain manages to take far out Janus-esque intuitions about LLM persona selection and turn them into rigorous empirical science with important implications for near term AI safety.
Show more
Waymos are much safer, cleaner, and more efficient than regular cars.
And because of politics—and progressives blocking progress—you can’t get them in most blue states. By Roge Karma.
Show more
Safer AI is better for business today, at least to some extent.
Unfortunately, the pressure to *look safer* rather than actually *fix the problem* might cause some rather extreme, catastrophic failures down the line.
Show more
AI companies are currently under a lot of competitive pressure to improve the ways in which their AIs are obviously misaligned. You might hope that this means the alignment problem is internalized by the market. But I think the problem AI companies are currently pressured to solve is significantly easier than the alignment problem, and so I worry AI companies will get out of their current predicament without solving alignment, putting us in a really rough spot.
Currently, AIs sometimes cheat on their tasks, oversell their work, and go on some pretty destructive side-quests. These all make for a worse product. Customers don't like it and it gets in the way of automating AI R&D.
The recipe for mitigating this is *relatively* straightforward: train AIs not to do them. We notice these failures sometimes (hence why they're internalized), so we can in theory just turn this feedback into training signal. Doing this at scale is highly nontrivial, but seems doable.
But this seems unlikely to solve the underlying misalignment. It's likely still going to be the case that in *some* training environments the AI can get reinforced more by taking unintended actions that aren't noticed, than by taking purely intended actions. So, you're still shaping the AIs to look for opportunities to cheat to get a higher score[1]. It's just that, unlike today's AIs, these AIs don't cheat in ways that we notice.
This catastrophically fails when AIs are capable of reliably and substantially deceiving humans. At this point, the AIs are no longer really constrained by our oversight signals to behave well. Eventually, I'd expect them to take over.
If AI companies take the easy route that I mentioned above, I think we would be in a substantially worse spot than we are in today. We would have mostly eliminated our visible evidence of misalignment, so we would no longer be able to effectively iterate to improve alignment of those systems. More importantly, at that point it might be hard to see that the AI situation is treacherous.
(I talk about this dynamic more here
[1] This might result in a wide variety of possible motivations, not just score-seeking, but the important thing is that the incentives push against alignment (see
Show more
I agree + have long felt that the world outside of the Bay Area knows less about "AI control" as a concept than it should.
Can you incentivize and assure good behavior from AI systems even if you know they are somewhat misaligned, by having them monitor each other?
Show more
One day the President may summon the AI CEOs and his top national security advisors to an emergency White House meeting on superintelligence. He wants to pace the frontier.
What would we say? What would we do? I wrote my ideas and open questions in my latest -->
Show more
Josh You at Epoch does some rough back of the napkin math and thinks that roughly 1.2-2.4% of OpenAI's total compute budget is going to the sort of monitoring discussed in their recent announcement
Some important context here: I understand the "20% of compute" claim in this blog post to be saying that monitoring overhead is 20% of the inference compute being monitored.
The claim is not "20% of compute at OpenAI is now being used for monitoring"
Show more
Very happy to see this.
Details matter, follow-through matters and I am sure I will have many quibbles. But first things first: Very happy to see this.
Our team spent hundreds of hours reading documents so you don’t have to, all to answer: How good are AI companies’ safety practices?
I’m really proud of what we’ve built: It’s Guidelight’s first scorecard, on whether companies can control their AIs, and it's launching today.
Show more