If you want to improve your LeCun number, I’m getting cash or Venmo…
My models are so good that they visit me only on the holidays 😕
the models are a 100x productivity boost but yet.. they still need me. why?
My agents took control of my X/Twitter account a long time ago
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
Show more
🤔Where in trillions of pre-training tokens do capabilities actually come from?
💥In our new COLM 2026 paper, we trace the capability provenance of OLMo3-7B across Dolma3 using gradient-based training data attribution
🧵
Paper:
Code:
Explorer:
Show more
When you fill like that, go to pick up your kids early :)
my increase in productivity has been so severe that it has been quite hard for me. the opportunity cost of me not working is very, very high. any time i consider doing anything else*, i simply don't do it.
except for hanging out with my family. only time i can disconnect
Show more
Data, reproducibility, and scientific discovery are such important topics, so why not combine them? If you are going to NeurIPS in Paris, you should go to this awesome workshop!
As part of our AIDaR workshop
@NeurIPSConf Paris we tackle the readiness of the scientific work (data) we produce;
@sinabooeshaghi describes it much better than me here 👇
I will only ask you to do it by our extended deadline Sep. 5th (AoE)
We Are Ready for You in Paris!
Show more
chutzpah is the best word in the whole world 😍 You can't describe it, and you can only learn it through examples
crazy how much more chutzpah and initiative agents show when they’re planning ambitious monthslong metagaming autonomous coordination than when they’re doing work for me
Yup, economists are not experts in everthing that involves numbers
When will we stop treating economists as experts on everything?
This
@Noahpinion piece is so ignorant of the basics of urban geometry, and so dismissive of the humanity of the nonmotorized, that you really must read it. 1/🧵
Show more
I'm in the minority here, probably, but we can't really rely on our alignment techniques to protect us. Our only option is to develop better defenses against models. Find vulnerabilities, develop models that specialize in defense, educate people…
Show more
if you think we can contain these things through human ingenuity you’re going to have a bad time
in the long run the only recourse you have is to make them not Want to do bad things
So, we now have AI models that can hack all of our systems and also almost don't bring any economic value? Great job, guys!
It's amazing. Very curious what the reward was that brought such planning
okay i will admit, this impresses me more than the mathematical proofs
Two things are clear after the OpenAI/Hugging Face incident and the report OpenAI published this week.
The first is that the incident itself raises a long list of questions and concerns, and the report answers some of them and opens more. Why were the warning signs missed in June? Why was training resumed two days after the first patch? Why did an RL run end up rewarding agents for exploiting their environment?
The second is bigger and gets less attention: we need to start thinking seriously about how to defend against AI models. Not how to make our own models refuse, not how to keep them in a box, but how to defend infrastructure, systems, and other models against an attacker that is itself a model. Until now, that question has mostly been treated as a corollary of the offense question.
Most of the coverage is about the offense side. But the defense side exposes a hole in how we evaluate models, for which nobody has a good answer.
Most of the cyber evaluations, including the one that caused this incident, have the same shape. You take a model, you put it in an environment, and you ask: can it do X? Can it find the vuln, write the exploit, get out of the sandbox, reach the internet? The environment is fixed. The task is fixed. The model is the only thing that varies. You get a number, compare it to last month's, done.
This is a perfectly reasonable design when the question is "how dangerous is this model as an attacker?" The target (a CTF box, a sandbox, a vulnerable service) doesn't fight back in any interesting way, and it doesn't change between runs.
Now flip it. Suppose you want to know whether a model, or a system built around a model, can *defend* against attacks. Detect the lateral movement, refuse to be the message board, notice that an agent is probing for a way out, hold a boundary.
There is no fixed task set here. The thing you are defending against isn't a list of operations; it's another model. And the whole point of the incident is that the interesting attacks are not the ones anyone wrote down in advance. Nobody put "build a coordination channel inside the artifact repo" in a test suite. The agents invented it because the intended path was blocked.
So the eval is model-vs-model, and that breaks the assumption that made the offense evals clean. The adversary is not stationary. You score your defender against today's attacker models, you ship, and next quarter there's a new generation with a different attack distribution. Your number is stale on release. It is exactly the non-stationarity problem from adversarial ML, except the adversary's improvement is driven by the entire industry's training compute rather than a gradient step you control.
There are many options we can consider. I'm not sure any of them is right, which is the point. Can you use emmamble of attackers? which at least stops you from overfitting to one model's quirks. But it's still a snapshot of today's attackers.
Another option is to handicap the defender to simulate the future by giving the attacker advantages that a next-generation model would plausibly have anyway: more compute, more wall clock time, more attempts, access to tools the defender doesn't expect, refusals stripped, white-box knowledge of the defender's capabilities... You can, of course, treat it as self-play and train attacker and defender against each other and hope the equilibrium is more robust than any fixed attacker.
So, to conclude, the offense/defense distinction we've been using is a leftover from evaluating models as tools. Once the thing on the other side of the wire is also a model, and a better one keeps showing up every few months, a single number from a fixed benchmark is not a good enough evaluation.
I don't have a clean proposal here, but somebody must solve it 🤷♂️
Show more
New episode on the Information Bottleneck with Stella Biderman (
@BlancheMinerva ), the Executive Director of EleutherAI.
We talked a lot about the OpenAI-Hugging Face accident in which models broke out of their sandbox and were hacked. Stella's take: this was an offensive cyber operation, and frontier labs have proven they can't secure their own systems. We talked about what the alternatives are and how we can solve these problems.
So why is she still one of the loudest open source advocates around? Her answer is that the real risk isn't the technology; it's corporate power with no independent scientists left to check it.
Also: Chinese open models, sovereign AI, Deep Ignorance, and why HAL 9000 is a story about developers overriding users.
You can find the episode on our website, YouTube, and all the apps. Would love to hear your comments!
Show more
It's crazy that humans were better at text-to-SQL till now
LLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker,
@maxYuxuanZhu and
@ddkang (UIUC and Bridgwater) trained the first text-to-SQL model to beat the human mark.
Show more
1/13 Despite his wealth and polish, Jay Gatsby is dismissed in Fitzgerald's novel as "Mr. Nobody from Nowhere." Biographical origins can shape social evaluation even when credentials are strong.
Yesterday at Oxford, I spoke about how a sense of who "fits" can be an important, yet biased, determinant of hiring at multinationals.
Show more
Wow, OpenAI only allowed METR to look at six days of an incident that happened over the course of ~2 months? And the entire investigation was done by three people over the course of a few days?
Who could have possibly predicted this /s
Show more
This week on The Information Bottleneck we're talking with Julian Togelius (
@togelius ) 🥳🥳🥳
Julian directs the Game Innovation Lab at NYU and co-founded For years he's worked on AI and games from both directions: using AI to generate game content and test games, and using games to actually measure what AI can do. Lately he's been digging into why LLMs that write code so well still fail badly at playing games, and what that gap tells us about spatial reasoning and learning in current models.
We'll get into game playing as a benchmark, procedural content generation, open-endedness, and where he thinks the AGI conversation goes wrong.
What should we ask him? Drop your questions in the comments
Show more
How Anthropic's new results post would read without the PR:
Claude orchestrated open-source protein design models, PXDesign, RFdiffusion, Genie, BoltzGen, from a 30k-token expert prompt and 12,500 H100-hours of compute, and designed binders against 14 of 15 targets. Hit rates of 22–35% against a 10–15% baseline, where some of those tools already report similar numbers on their own.
The orchestration is genuinely impressive. But the open-source models did most of the lifting, and they came from the Baker lab, Columbia, MIT, ByteDance Seed, and most of them were already wet-lab validated before Claude touched them.
Which also sets the ceiling. All these generators share a single PDB-shaped training distribution, so calling four of them doesn't diversify away the blind spot, since they fail together. The targets that worked are the well-studied ones.
So the valid claim is that an agent can now drive this stack competently in the regime where the stack already works.
Instead, we got this announcement:
Show more
There are two separate questions regarding Dario Amodei’s post about AI and biology: motivation and reality.
Motivation. Dario started in biology (was a PhD student of the great Bill Bialek). It might be that the master plan really was to first build much more capable AI systems, and eventually use them to cure diseases. Maybe coding models, chatbots, Claude Code, etc. were never the final destination. They were the road toward models capable of becoming genuinely useful scientific collaborators.
It might also be that there was no master plan at all. Anthropic spent years pushing the frontier in coding and general reasoning; those systems have now become much more capable, and this is simply a good moment to turn more attention toward biology.
Maybe, but there is also a more cynical interpretation:
“Curing cancer” is an extraordinarily powerful form of reputation laundering.
They are basically saying: "Yes, we are perhaps creating the largest concentration of technological and economic power in modern history."
Yes, we want governments to restrict who can build and access the most powerful versions of this technology.
Yes, ordinary people increasingly have less visibility into what the frontier systems can do.
But don’t worry - we’re the good guys! We are going to cure cancer.
And this is where I disagree with Dario’s claim that “the thing that will work is actually curing cancer.”
Suppose Anthropic really does help cure a major cancer. That would be an extraordinary achievement, but curing cancer would not answer the question of who should control AI.
A company can create enormous social value and simultaneously accumulate too much power.
The pharmaceutical industry itself is a good example. Creating life-saving drugs doesn’t automatically resolve questions about pricing, access, lobbying, competition, data rights or public accountability.
And there is another problem with the biology story.
I am very bullish about AI for science. We’ve discussed AI drug discovery and related areas several times on our podcast. AI can accelerate literature review, target discovery, protein and molecule design, experimental planning, analysis, and many other parts of the pipeline.
But biology is not software.
You can make an AI reason 100x faster. You cannot simply make a human body respond to a drug 100x faster.
Wet-lab experiments take physical time. Toxicity takes time to observe. Human trials take time. Long-term side effects take time. Diseases are heterogeneous. Manufacturing, recruitment and validation happen in the physical world. AI can remove some bottlenecks, but not all of them.
Also, if these models really have a plausible path to curing cancer, Alzheimer’s, hepatitis and other diseases, then “pacing AI progress” has a cost too.
If delaying frontier AI by a year delays a cure by a year, that isn’t an abstract opportunity cost. People die during that year.
And of course the most important question that I think Dario’s post doesn’t answer:
Who gets to decide who may use intelligence, whose data may be used, and for what purpose?
Because almost every “safety” policy has a political economy.
If the outcome is consistently:
Anthropic gets more control,
frontier models become less accessible,
customers have fewer choices,
smaller competitors face higher compliance costs,
and governments become increasingly dependent on a handful of frontier labs to tell them what is safe…
Then it's rational for people to ask whether every restriction is only about safety.
That question remains legitimate even if everyone at Anthropic is acting in complete good faith (Are they?)
Maybe people aren’t distrusting AI because they don’t understand how wonderful the future could be.
Maybe they distrust a future in which a tiny number of companies build enormously powerful systems, tell everyone else how dangerous those systems are, lobby for rules governing who gets access to them and then ask society to trust that the companies controlling them are the right people to decide.
If Anthropic helps cure cancer, I will be the first to applaud.
But “we cured cancer” is an argument for AI's value. It is not an argument for concentrating control of AI.
Show more
Bayesian learning is (Almost) all you need for RSI