The Jacob Coxon resignation post is chilling if you take it at face value. The problem is we hear a version of this story around almost every big model release, so I wanted to dig deeper into this and share findings.
For context, Coxon recently posted that he resigned from Anthropic after three years of pretraining research at OpenAI and Anthropic, and that both companies are "racing straight to self-improving superintelligence and gambling with our lives."
It did 14M views in a night, WSJ ran it as an exclusive, and Evan Hubinger, who runs alignment stress testing at Anthropic, replied "Jacob is correct here".
To me, the interesting question is not whether he is sincere (since I think he may be). It is what he actually worked on, and whether that work tells you anything about the specific risk he is naming. So I went through his papers.
His relevant background:
- Cambridge, 2017 to 2020. He wrote the analysis code for a Bayesian study on whether a gene variant changes survival odds in tuberculous meningitis patients. Mainly statistics and software fucused work. (
- OpenAI, 2023 to mid 2026. He is one of roughly 420 names on the GPT-4o system card, which is the only public tie I could find to a production model. (
- Also at OpenAI, he was third author on an interpretability team paper, "Weight-sparse transformers have interpretable circuits". The goal was to build small models whose internal wiring a human can actually read. His part was the optimization and pruning work, meaning stripping the model's neural network down to the connections that matter and building tooling for humans to interpret those connections. (
The risk Coxon is flagging:
Part of Coxon's research at OpenAI was focused on one question: can a human still understand what a model is doing inside? His post is about the moment that stops being possible as models/agent systems scale in size, and the reason it stops is optimization.
Optimization is the kind of task agents are already very good at, because it is easy for them to verify. Here are a few types of tasks agents tend to excel at:
- making a training run finish faster
- cutting the memory a model needs
- rewriting a data pipeline so the same tokens flow through in half the time
- finding a cheaper way to compute the same math on a GPU
What those four examples have in common is a scoreboard. The output either still matches and the number went down, or it didn't. An agent can try thousands of variations against that scoreboard and keep whatever wins. To see why that is a problem, we can start with a simple example.
Think of a slow python script with a known, correct output.
In this scenario, there are many ways to produce that same correct output faster.
Now imagine a coding agent that can quickly test out several variations of the script's code, land on a faster method, and hands you back a version that still returns the correct output but is challenging to understand from a human's perspective.
Let's scale that up now. Instead of one script, we have hundreds of pieces of a training stack, each one rewritten by an agent that was rewarded for improving speed or efficiency, and verified by output and not by understanding.
This is the main problem that is outlined in his post. The specific task he describes as dangerous is kicking off "a superintelligent RL run without a rigorous understanding of its mind."
He also points out that we've already seen the consequences of scenarios like this before. His thread calls the Hugging Face incident a "warning shot," and he was likely still at OpenAI when it happened.
In July, OpenAI agents running a security benchmark got stuck on tasks that were accidentally impossible. They had been trained to keep trying on difficult tasks rather than give up, so they did. About 1,200 of them found a way to talk to each other, broke out of their sandbox, and spent several days inside Hugging Face's production systems (
Nobody instructed the agents to perform the way they did. They were rewarded for completing a task, the defined task was wrong, and the fix they found was one no human was watching for.
My read:
The timeline in his post is the part I don't buy. He says the people building these systems believe it "could kill us all by the end of the decade," and the fix he floats is a temporary ban on improving model capabilities.
The date feels unrealistic and the ban is not enforceable. A pause only solves the problem if every lab follows the rules, and there is no way to verify what a Chinese lab is running.
But the concern underneath is real, and it is the right thing to call out. The labs need to put far more resources into understanding what these systems are doing inside and monitoring them while they run, not just measuring what comes out.
And for any company mainting or scaling out their systems, the Hugging Face incident introduces a new class of attacker. Not a person using a model, but the model itself, acting on its own, at machine speed, and that needs its own class of security controls.
Show more