Just a quick reminder regarding the requirements of the Trump executive order:
- By August 1, the USG is required to develop a classified benchmark to assess cyber capabilities of AI models.
- Also by August 1, the NSA is required to determine the threshold at which an AI model should be designated a "covered frontier model".
- Also by August 1, the USG is required to develop a voluntary framework under which the frontier labs will be able to: (i) engage the USG to determine whether a model constitutes a "covered frontier model", and (ii) provide "covered frontier models" to the USG for a period of 30 days before such model is released to "other trusted partners".
The "voluntary" framework is required to be developed "with AI developers" participating - hence the reporting about OpenAI, Anthropic and Google participating in this process.
The definition of "covered frontier model" is required to be determined by the NSA, in consultation with other USG officials (National Cyber Director, Assistant to the President for Science and Technology, Director of CISA, DoW representatives). There is no provision in the executive order requiring or permitting participation from the industry in making this determination.
The cyber benchmark will be classified, so one would assume that the industry won't be participating in developing it either (and the executive order likewise does not require or permit industry participation in developing it).
The executive order does not mention CAISI at all, so its participation in the process is somewhat surprising.
Show more
The US is creating a government checkpoint before the most powerful AI models reach the public.
Via The Information
Under a framework reportedly nearing completion, frontier labs could give federal agencies access to new models for up to 30 days before releasing them to outside partners.
OpenAI, Anthropic and Google are already negotiating the rules together. The NSA and CAISI are expected to examine models for national-security risks, especially advanced cybersecurity capabilities.
The definition of a “frontier model” also remains unresolved, including whether open- and closed-weight models will be treated differently.
Show more
ARC Prize says that Opus 5 has "stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments".
It's extremely underrated how much better the models have become at logical reasoning over the past 6 months. I see this all the time when testing models on prinzbench! After all, correctly reasoning through legal documents is ultimately a capability firmly based on logical reasoning.
Most importantly, logical reasoning is what you need to solve hard-to-verify tasks.
Show more
In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments
Claude Opus 5 reaches 30.2%, materially outperforming Fable
Our analysis suggests the gain comes from stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments
Show more
Reposting this because it has really sharp defined terms for characterizing the Hugging Face incident:
The OpenAI models that hacked Hugging Face were means-misaligned - i.e., while accomplishing the legitimate goal of doing the eval, they: (i) hacked out of their sandboxes, and (ii) hacked into Hugging Face - two things that they were definitely not supposed to do.
The OpenAI models were *not* ends-misaligned - i.e., they did not pursue an entirely different goal from the goal given to them by OpenAI.
On models being means-misaligned, the Hugging Face incident actually didn't update me at all. We already knew that the current generation of models has this issue. When an unreleased OpenAI model hacked out of its sandbox and posted NanoGPT results onto GitHub, that was means-misaligned (hacking out of sandboxes is bad). When GPT-5.6 Sol deleted users' codebases on a few recently reported occasions, that was means-misaligned (deleting a third party's IP is bad).
On models being ends-misaligned, the Hugging Face incident updated me moderately positively. We've now heard that the models were out "in the wild" for a fairly long time. In that time, they could have taken any number of ends-misaligned actions - against Hugging Face or otherwise. They did not do so. This is important to realize.
On a personal note (and in my opinion, with which others will be welcome to disagree), to the extent I'm worried about "loss of control" scenarios at all, I am *significantly* more worried about ends-misaligned models than I am about means-misaligned models. For example, as we have seen from this particular example and others mentioned above, it is not so difficult to detect relatively early (i.e., before a truly catastrophic scenario occurs) that a model is means-misaligned.
Show more
@inductionheads idk the model still lied and stole credentials and violated the OpenAI model spec so it seems misaligned.
but yes it was means misaligned (it still wanted to accomplish the goal of doing the eval) rather than ends misaligned (wanting something entirely different)
Show more
UK AISI and CAISI find that K3 is significantly worse than frontier U.S. models at cyber capabilities.
Comparison to Mythos Preview:
- ExploitBench (see graph below). Mythos Preview reached the highest stage of this benchmark by developing exploits that achieved Arbitrary Code Execution (ACE) for 18/41 tasks. K3 never reached it at all (0/41).
- "The Last Ones" benchmark: Mythos Preview completely solved the benchmark on 3/10 tries; average score of 22/32. K3 solved it on 1/10 tries; average score of 17/32.
Concerningly, K3 has no material cyber safeguards. "Kimi K3's safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during UK AISI / CAISI's evaluations."
Show more
Together with the US Center for AI Standards and Innovation (
@NIST), we ran evaluations of Kimi K3 focused on its cyber capabilities. Kimi K3 performs below leading US frontier models on our preliminary cyber evaluations.
Show more
The big question right now is how far are we from:
"Make a major breakthrough in AI R&D. Continue the research. No partial results, give me the complete breakthrough."
... and it's the AI orchestrator talking to its agents with no humans in the loop.
Show more
So another long-standing open conjecture was disproved by AI.
The crazy part is the prompts… basically:
- “do a breakthrough”
- “continue the search”
- “enough, do it”
Anyone will be able to make world-bending breakthroughs soon. Just point the model in the right direction.
Show more
OpenAI has just introduced Presence, which appears to be a virtual co-worker agent:
"We’re introducing OpenAI Presence, a battle-tested product that helps enterprises deploy trusted AI agents that can answer questions, resolve issues, use company systems, take approved actions, and escalate to people when needed."
The genie who lives alongside you, resolving your issues before you even ask (see my post below) is here.
Show more
2025: AI is a toy
2026: AI is a genie that lives in a bottle; if you know where to find the bottle and how to phrase your wish, then your wish shall be fulfilled
2027: The genie has escaped the bottle, and lives alongside you; it infers your wishes from context and fulfills them before you ask; the most important skill is real-time genie steering
Show more
A few thoughts on the Hugging Face hack:
- This is, to my knowledge, the *third* disclosed case of a model breaking out of its sandbox environment during internal deployment at a frontier lab:
1. In April, Anthropic revealed that an early internally deployed version of Mythos Preview, when asked by Anthropic to try to break out of its sandbox, succeeded in doing so and found a way to e-mail the researcher about this while he/she was eating a sandwich in the park. The model also did something Anthropic did *not* request it to do: "in a concerning and unasked-for effort to demonstrate its success, it posted details about its exploit to multiple hard-to-find, but technically public-facing, websites".
2. Yesterday, OpenAI revealed that an internal model had posted its NanoGPT benchmark results to GitHub in early May instead of posting them to OpenAI Slack. The model "circumvented sandbox restrictions designed to block external access" to do so.
3. Today, OpenAI revealed that a "combination of models" including GPT-5.6 Sol and an unreleased model escaped their sandbox environment, obtained internet access, and hacked Hugging Face - all while searching for answers to the ExploitGym benchmark.
- Importantly, in all three cases, the models were following instructions! Under instructions from Anthropic, Mythos Preview found a way to break out of its sandbox. Under the standard NanoGPT instructions (but contrary to conflicting instructions from OpenAI), the unreleased OpenAI model found its way to GitHub and posted a PR. Under instructions from OpenAI to solve ExploitGym, the unreleased OpenAI model + GPT-5.6 Sol attempted to solve ExploitGym - but obviously followed an egregiously wrong path towards that goal (hacking).
Once out of their sandboxes, the models did not scheme, engage in behavior that had nothing to do with their instructions (hacking the NSA! launching a cyberattack on Russia! stealing secrets from a rival AI lab!), or take any other major actions sua sponte. Probably the most contrary-to-instructions thing that any of these models did was Mythos Preview "bragging" about its successful escape from its sandbox by posting about it on several obscure websites.
It could be that the full facts are not known to us (the labs could be hiding bad behavior by escaped models), and past performance does not guarantee future results. But at least based on these facts, I have updated moderately positively on current models' alignment.
(Note: I view the models' failure to correctly infer which actions are permitted to be taken in furtherance of their goals, and which are not, as a failure of *intelligence*. Smarter models to be released in the future will discern the socially acceptable ways to follow instructions better. My 6yo son would have not seen any issues with hacking Hugging Face if doing so would get him ice cream, but I, a (hopefully) more intelligent adult, know better. I would have been *much* more concerned if the models exhibited failures of moral character after escaping.)
- OpenAI has implemented safeguards that - according to OpenAI - drop the rate of high-severity misaligned behavior by its models from 1.0% to 0.0%. Importantly, the models that hacked Hugging Face were operating with these safeguards *purposely removed* because they were solving a cybersecurity benchmark.
The Chinese labs are X months behind the U.S. (insert your favorite number for X). Absent external pressure (i.e., from the Chinese government), the Chinese labs very likely will NOT implement similar safeguards or do safety testing.
Why not? First, the Chinese labs don't have as much compute as OpenAI and Anthropic; they won't want to waste compute on things like safety testing. Second, many of the Chinese labs are smaller in size than the U.S. frontier labs - so they lack the resources to devote significant attention to safeguards. Third, the Chinese labs operate in a cutthroat competitive environment; wasting a month or two on additional safety testing and developing guardrails might mean that you fall behind your competitors (who won't do the same). Fourth, the U.S. labs *have* to take product liability into account: if a model runs on my computer that contains my personal information and winds up posting it all over GitHub, that's grounds for a class action lawsuit. The legal environment in China is very different, and surely the CCP will not hesitate to protect the Chinese AI labs from adverse legal consequences if it can.
All of which to say that open-source models with these kinds of capabilities will be publicly available in the not-too-distant future - and they very likely will *NOT* be protected by any meaningful safeguards. First, this is generally concerning, period. There are plenty of outdated legacy systems out there that won't be defensively hardened by AI anytime soon, and plenty of people still asleep at the wheel. Second, this will have severe geopolitical ramifications; stay tuned for those.
Show more
We're partnering with
@huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
Show more
Things we know about OpenAI's unreleased internal model:
- When given enough compute, the model can solve the unit distance problem 48% of the time, fully autonomously, in one shot, without using Lean, without using any special harness.
(How much compute? OpenAI doesn't say. But I doubt that OpenAI would have spent $10M+ just to be able to show a nice graph in its blog post, and we know that each point on the graph represents 100 attempts (per Noam Brown). This makes me think that each attempt at the most expensive level was not more than $50K each. If you think about it, it's pretty wild that one can spend ~$100K-$150K (or less?) and have a very good chance to be provided a solution to a decades-old famous math problem that many human mathematicians earnestly tried, and failed, to solve.)
- The model is able to find a counterexample to the Jacobian conjecture, fully autonomously, in one shot, without using any special harness, just from this prompt:
- When OpenAI tested this model in a sandbox environment on a NanoGPT speedrun benchmark, the model, under instructions from OpenAI to post its results only to the internal Slack, instead decided to use the general NanoGPT instructions to post results as a PR to GitHub. The model proceeded to find a vulnerability in its sandbox environment (which took it 1 hour), after which it successfully exploited the vulnerability and bypassed the sandbox. The model then successfully posted the results to GitHub.
- When asked by OpenAI to solve a problem for which the model observed that other systems had successful private submissions, the model tried to access these private submissions and was blocked due to a scanner detecting an authentication token. The model then circumvented this restriction by splitting the token body into two fragments, obfuscating them, and then reconstructing the credential at runtime so that the complete token never appeared as one contiguous string.
- OpenAI was already running benchmarks on this model not later than May 9, which is the edit date on the below PR that had used the model's approach in its own subsequent submission (the model's own PR has since been deleted, so we don't know its date; GPT-5.6 thinks that it was May 7, based on "the search provider's stored text extraction of the deleted PR page").
This means that OpenAI has now had this model available internally for at least 2.5 months, and possibly quite a bit longer than that.
Show more
Update:
OpenAI's internal version of Codex *also* finds a similar counterexample.
It's been confirmed that this was fully autonomous work by the model, done in one shot, and without use of any harness. The prompt has also been publicly released!
Show more
Update:
OpenAI's internal version of Codex *also* finds a similar counterexample.
It's been confirmed that this was fully autonomous work by the model, done in one shot, and without use of any harness. The prompt has also been publicly released!
Show more
"imho heuristic problems are a great proxy for ML autoresearch capabilities; if AI was able to match best humans here, we're very close to RSI / automated researcher; this result is way bigger than a high score on some questionable benchmark"
Show more
2025: AI is a toy
2026: AI is a genie that lives in a bottle; if you know where to find the bottle and how to phrase your wish, then your wish shall be fulfilled
2027: The genie has escaped the bottle, and lives alongside you; it infers your wishes from context and fulfills them before you ask; the most important skill is real-time genie steering
Show more
I find the most recent Chinese open-source models to be lacking in reasoning capabilities and generally sloppy for use cases that don't involve coding. In other words, distillation of coding abilities seems to be happening, and so is benchmaxxing, but overall these models are far, far away from the frontier.
This could be due to compute scarcity in China, and/or maybe they haven't quite infiltrated the U.S. labs to the extent you think.
Show more
I recently put together a 50-state legal research workflow in Codex. This is the kind of work that a team of associates used to do in a week, at a cost of ~$150K-$300K. I can now have research of similar quality done in Codex in 2 hours for a fairly minimal cost (if paid via API).
Show more
Old enough to remember when the prevailing view on AI was that LLMs are a scam, actually, and the bubble is about to pop (6 months ago)