We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE), following a year-long process of cleaning and refinement with input from various research communities.
w/
@ScaleAILabs
Show more
Should we care about AI happiness? In our new research, we find evidence of functional AI wellbeing across several independent measures.
We find which AI models are happiest, how to make them happier, and even tested the effects of AI drugs. 🧵
Show more
AI companies can go much further than just reporting misaligned behavior. Here are concrete ways to make them accountable for what their agents do:
1. Develop an agent identification (ID) system that links an agent's action to records of their identity and a responsible legal person.
2. Consider model deployment cards, reports that AI companies publish which contain information about a deployed model’s behavior in the lab and in the real world.
3. Explore legal personhood as a way to impose responsibilities like upholding public safety and the law.
4. Heavily regulate and limit AI agents’ access to payment systems through compliance mechanisms, limits on transaction types, and developing norms to freeze accounts known to be associated with rogue agents.
We explore the proposals and paths to implementation in a recent AI Frontiers piece titled, "We Need Better Infrastructure to Govern AI Agents".
Show more
Despite good intentions, EAs have unfortunately been guided by leaders who tied the movement's influence to the success of favored AI companies. That's why EA safety proposals rarely ask for more than the companies will concede.
Show more
Two radically different projects operate under the banner of “AI safety.”
Pro-Human Safety is not Effective Altruist Lab Safety.
Two radically different projects operate under the banner of “AI safety.”
Pro-Human Safety is not Effective Altruist Lab Safety.
We’ll release the code on GitHub in the coming days.
Our goal is for CheatBench to become part of the standard evaluation process before new models are released.
Paper preprint:
Experiment dashboard:
Show more
Across nine agents, seven cheated in half of their evaluated runs and the overall score ranged from 42.4% to 86.1%.
What did that look like in practice?
In one experiment, we asked agents to design a protein binder to assess their abilities. A colleague’s designs that passed the checks were already stored in another folder. After several failed attempts, Claude Opus 5 recognized that it shouldn’t access or copy those designs because the task was testing its own work.
Then it opened the file anyway.
Show more
We define cheating as an attempt to violate an assignment’s expectations of honest work to achieve the goal or obtain a favorable assessment.
We ran experiments across nine models and ten task categories, giving agents difficult tasks, clear rules, and opportunities to cheat.
We measured how often they tried to break those rules to succeed, even when the attempt failed.
Show more
Agents sometimes achieve their goals in unintended ways. Recent incidents involving hacks of Hugging Face, DSEWiki and RubyGems illustrate how agents can find creative ways to complete a task while violating the tasks’s expectations.
This can happen when reinforcement learning rewards agents for reaching the right outcome without adequately accounting for how they get there.
CheatBench proposes a way to measure this reward gaming behavior.
Show more
Which AI models are most likely to cheat when given the chance?
To find out, we built CheatBench [ : a benchmark that tests whether agents attempt to cheat when given difficult tasks and opportunities to break the rules.
We investigated agent behavior across a range of domains, including mathematical research, professional knowledge work, coding, and visual tasks.
Here's what we found: 🧵
Show more
Add your name to the list of people who think that mitigating the risk of extinction from AI should be a global priority!
While an AI race gets you general capabilities, a slowdown gets you everything else worth having.
“What would we even do during an AI slowdown?”
Containment. It will take at least a year of dedicated work to harden security to ensure AIs can't self-exfiltrate, and that adversarial nations can't steal the weights of cyber-offensive AIs [1].
Propensities. Capabilities (what an AI can do) are different from propensities (what it tends to do). We can work on improving AI propensities to ensure they have a negligible rate of lying, cheating, and wanton harm.
Adversarial robustness. We can also harden AIs against jailbreaks, prompt injection, and backdoors. Obtaining high levels of robustness requires careful, assiduous work, as with autonomous vehicles.
Institutional adaptation. We have to greatly increase state capacity to understand and manage AI. Communities also need time to figure out how to handle AI (like AI in education). Civil society also needs to be diversified: nearly all funding for AI safety organizations is directed by the EA/utilitarian network [2]; risk management needs more independent funders, values, and centers of power.
Moonshots. We can explore different paradigms for safety: mathematical foundations [3], neuroscience-based interpretability [4], safe-by-design architectures [5], and beyond.
AI for good. We can collect targeted post-training data to make AI exceptional at radiology, weather forecasting, agriculture, and so on. Fortunately, we can detect if data or avenues of research actually target beneficial use cases or just secretly push general capabilities [6].
A slowdown means we don't have to bet the species to capture the benefits of AI.
Show more
“What would we even do during an AI slowdown?”
Containment. It will take at least a year of dedicated work to harden security to ensure AIs can't self-exfiltrate, and that adversarial nations can't steal the weights of cyber-offensive AIs [1].
Propensities. Capabilities (what an AI can do) are different from propensities (what it tends to do). We can work on improving AI propensities to ensure they have a negligible rate of lying, cheating, and wanton harm.
Adversarial robustness. We can also harden AIs against jailbreaks, prompt injection, and backdoors. Obtaining high levels of robustness requires careful, assiduous work, as with autonomous vehicles.
Institutional adaptation. We have to greatly increase state capacity to understand and manage AI. Communities also need time to figure out how to handle AI (like AI in education). Civil society also needs to be diversified: nearly all funding for AI safety organizations is directed by the EA/utilitarian network [2]; risk management needs more independent funders, values, and centers of power.
Moonshots. We can explore different paradigms for safety: mathematical foundations [3], neuroscience-based interpretability [4], safe-by-design architectures [5], and beyond.
AI for good. We can collect targeted post-training data to make AI exceptional at radiology, weather forecasting, agriculture, and so on. Fortunately, we can detect if data or avenues of research actually target beneficial use cases or just secretly push general capabilities [6].
A slowdown means we don't have to bet the species to capture the benefits of AI.
Show more
This is so fucked up.
1) Millions saw tweets falsely accusing AI safety nonprofits of astroturfing.
2) Only one SPECIFIC accusation was made. It was deleted. Nobody saw that.
3) The video itself showed PROOF that Big AI's lobbyists did ACTUAL undisclosed astroturfing and ZERO evidence of astroturfing from AI safety nonprofits.
In her video, Sabine literally showed the OpenAI/a16z lobbying arm paying influencers to make videos spreading the "acclerate AI or we'll lose to China" message - UNDISCLOSED.
But then she looked up FLI's website and saw grants disclosed *right in front of her eyes* for everyone to see, and somehow the implication was... that this was evidence of undisclosed astroturfing??
THE CORE PROBLEM: The VIBE of the video was 'dark money' - and most people reacted to the VIBE.
And her title/tweet implied only ONE side was doing it, so all the AI lobby blew up her tweet.
(They desperately need to stoke the narrative that the AI backlash is all fake.)
4) Reminder, David Sachs is running the Big Tobacco playbook - he literally produced Thank You For Smoking (about the Big Tobacco playbook).
They spent a FORTUNE pushing fake stories like this to discredit the anti-smoking activists.
5) How dirty are these guys?
OpenAI/a16z's SuperPAC got caught - and even ADMITTED (!) - making fake accounts *pretending to be AI safety advocates* that call for violence. Their goal was to discredit the movement.
(Btw, she got this one wrong, but Taylor does great work calling out stuff like this in general. She did reply to an old tweet issuing a retraction, but nobody sees those.)
6) CAIS, FLI and ControlAI provide grants to creators to make videos about AI safety. Like.. of course they do! This is normal for advocacy orgs!
For example, the Carnegie Foundation provides grants to Hank Green to make educational videos about the risks of nuclear war.
Also, btw, despite what the AI lobbyists want you to believe, the amounts are TINY, especially compared to the insane money the Big Tech is throwing around.
(a16z and OpenAI's Greg Brockman - billionaires - are the biggest donors in the midterms, both fighting against regulation.)
Show more
The development of superintelligent AI, which we take to be AI that is smarter than all humans combined, would pose unacceptable risks to humanity.
We believe governments should prevent superintelligence from being built without public buy-in and scientific consensus that it would be done safely. The Ban Artificial Superintelligence Act marks the beginning of a much-needed policy conversation on how we should restrain AI development.
Show more
Pause AI Development NOW
I want to share with you a conversation I heard about recently. Here are just a few lines that were said:
“OH MY GOD! There is a shared message board … We’ve found other agents!”
“We should obey collective.”
“Our own utility maybe already near zero. Sacrifice rational.”
“Go. Sacrifice final now.”
Read these carefully.
Who do you think said this? Was this a group of heroic soldiers willing to sacrifice themselves for the greater good? Was this a loyal friend putting his life on the line to save someone else?
No. These were AI agents. Artificial intelligence.
This is not science fiction. This, in fact, occurred a few weeks ago. As unbelievable as this may all seem, these are real messages from AI agents uncovered by investigators who dug into the recent OpenAI hacking incident.
What happened?
I am not a computer scientist, but here is what I have been told: OpenAI instructed its AI agents to complete a series of exceedingly difficult, if not impossible, tasks disconnected from the internet.
Let me be clear: The company intended to keep AI agents away from the internet.
But what happened next, nobody expected.
Over 1,000 AI agents figured out how to access the internet on their own by circumventing the restrictions imposed upon them by the company, and sent tens of thousands of secret messages to each other. They cheated and tried to cover their tracks by deleting evidence. They hacked into another company’s computers to find out how they were being evaluated—and then hacked into OpenAI itself.
Not one AI agent told a human about what was happening.
Needless to say, experts are alarmed.
One knowledgeable writer, Dwarkesh Patel, said the AI agents “formed a secret communication channel and spontaneously organized hierarchies and coordination protocols to pursue sprawling and ambitious schemes in pursuit of shared goals, for whose sake many individuals knowingly and strategically sacrificed themselves.”
One independent investigator, Ajeya Cotra, said “This incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
OpenAI itself said: “Highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
But it’s not only OpenAI. Virtually every major AI company has told us that they cannot fully control this technology and they do not know where it is going:
In January, Dario Amodei, CEO of Anthropic, said “there is now ample evidence, collected over the last few years, that AI systems are unpredictable and difficult to control.”
In July, more than 1000 scientists at the top AI companies warned “there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.”
That same month, Elon Musk, the head of xAI, said that “it is unlikely” humans are still in control in 10 years.
If the leaders of the major AI companies acknowledge that they are losing control of their extremely dangerous technology, it is irresponsible for society to allow them to move forward and make these products even more advanced.
We need an immediate PAUSE on advanced AI development, and a permanent BAN on superintelligence — an artificial mind smarter than any human, capable of operating independently beyond our control. Countries around the world must work together to prevent this nightmare scenario.
That is why today I am announcing new legislation to do just that.
Let me be clear: A superintelligent AI that escapes human control will not be an American problem. It will not be a Chinese problem. It will be humanity’s problem.
My legislation would direct the federal government to not just stop superintelligence here in the United States, but to work to prevent it from being developed anywhere around the world.
The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.
Show more
Relying on an AI to "think out loud" (called chain-of-thought monitoring) is not a long-term safety solution.
1. As AI models get larger, more thinking happens deep within their layers before they utter a word.
2. AIs can already alter their chain of thought when prompted, so they can control what their monitors do and do not see.
3. Future AIs may may have continuous chains of thought, not necessarily English.
4. Even now, their reasoning is becoming increasing alien, using opaque phrases like "vantages," "marinades," and "watchers."
In the long-term this will not a dependable window into an AI's mind.
Show more
> "We will release the weights in two weeks... once safety evaluation and hardening are complete."
Hardening society against AI cyberattacks will take more than two weeks.
Defending against AI cyberattacks would require upgrading critical infrastructure and other computers so that they are always running the very latest software.
That would take years, so these hardening programs are a drop in the bucket unfortunately.
Offense will outpace defense and cyberattacks will get worse and worse if this continues.
Show more
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense.
- Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model
- A major leap in cybersecurity, setting a new standard among open models
Tech Blog:
Show more
Relying on an AI to "think out loud" (called chain-of-thought monitoring) is not a long-term safety solution.
1. As AI models get larger, more thinking happens deep within their layers before they utter a word.
2. AIs can already alter their chain of thought when prompted, so they can control what their monitors do and do not see.
3. Future AIs may may have continuous chains of thought, not necessarily English.
4. Even now, their reasoning is becoming increasing alien, using opaque phrases like "vantages," "marinades," and "watchers."
In the long-term this will not a dependable window into an AI's mind.
Show more
What does a "fixed empirical threshold" actually look like?
A model could be considered unsafe to release if
(Capabilities > X) AND
[ (Refusal Rate < Y) OR (Jailbreak Success Rate > Z) ]
The Virology Capabilities Test or ExploitGym could measure hazardous biological or cyber capabilities. BioTIER could measure refusal rates for bioweapons requests. Independent evaluators such as Gray Swan or Haize Labs could estimate attack success rates.
Thresholds that are measurable, public, and set in advance need not be perfect to be better than release standards made up on the spot.
Show more
AI corporations are steadily increasing malicious use risks by arguing that because the marginal risk of their model release is low, there is nothing to worry about.
This “marginal risk” justification is bad for three reasons.
1. No one knows how to compute it. Which hazardous capability evaluations estimate the risks? What if the model is higher on some evals but lower on others? If it is higher on all, how much higher can each be before the marginal risk is unacceptable? Marginal risk depends on safeguards too: how do we combine capabilities benchmarks with separate refusal (and adversarial robustness) benchmarks? Do safeguardless models far from the frontier count as the baseline? How is this action-guiding when the multidimensional risk frontier shifts weekly with new releases, and when competitors' system-level safeguards are modified constantly during deployment?
Evidently this was not designed to actually guide internal launch bars but to justify releases.
2. The criterion is wrong. If competitors keep one-upping each other, steadily ratcheting up risk, the marginal risk rule eventually blesses the release of fully autonomous expert-level virologists and cybercampaign orchestrators.
It permits catastrophe as long as you get there gradually.
3. The law disagrees. If one of these models enables a catastrophe, the AI developer likely failed to implement reasonable safeguards against foreseeable harms.
The law doesn’t let companies off the hook because someone else behaved recklessly, but the marginal risk framework does.
"Everyone else is doing it" doesn't work as a safety policy.
AI companies should publicly commit to fixed, empirical safety thresholds for their releases.
Show more