Register and share your invite link to earn from video plays and referrals.

Nathan Lambert
@natolambert
Open model research @ something new. Prev. co-led Olmo at Ai2. Writes @interconnectsai, wrote
943 Following    102K Followers
Frontier: pacing Pulling up the ladder: accelerating
Claude Opus 5.5 will be the first Opus meant to fall back to a less capable model for "a small set of capabilities related to the development of frontier LLMs, such as kernel development.."
Show more
Whenever I see a model top this chart, it's normally a bit of a red flag. Would love to be proven wrong!
As expected, unfortunately the only bad thing about the new Opus is it's a token-gobbling monster. Uses the most tokens per Intelligence Index task of any model benchmarked. The price cut makes more sense here!
Show more
The biggest disagreement JSD and I had in this podcast was on the impact of distillation. In our research for Interconnects, @xeophon and I agree that there's no hard evidence that distillation is a massive impact for the Chinese labs. At the same time, the gossip mill in SF has been doubling down on the Chinese labs getting massive gains from distilltion. The argument where distillation is a huge impact is something along the lines of distilled traces go into mid training and make RL work far more easily. My argument, that distillation is a 1-2 month pull ahead in capabilities closer to the frontier, is that Chinese labs already have sufficiently strong models, where this mid-training setup can be done on their own models, and most of the capabilities gains are from scaling RL environment training, which doesn't link cleanly to the distillation data pathway. We feel like the recent RL dashboard from @XiaomiMiMo supports this claim. A lot of it comes down to a gut call on if you can believe it that the Chinese labs are really great at building LLMs, maybe even better than focusing than OpenAI/Ant etc, as their current ambitions are a bit narrower (catching up), rather than transformative products/inventions
Show more
New podcast with @datagenproc of @EpochAIResearch digging into the open questions determining the future of frontier AI! We cover: 00:00 Predictions for RSI 18:15 The role of robotics in an AI acceleration 24:20 How far behind are Chinese models? 27:39 Does distillation explain the gap? 40:58 What Chinese job postings reveal about their labs 48:13 Are open or closed models safer? 58:10 How Epoch AI ticks 1:00:55 What a frontier post-training recipe looks like He's one of the people who gives the best feedback on my writing, so I was stoked to have him on.
Show more
New podcast with @datagenproc of @EpochAIResearch digging into the open questions determining the future of frontier AI! We cover: 00:00 Predictions for RSI 18:15 The role of robotics in an AI acceleration 24:20 How far behind are Chinese models? 27:39 Does distillation explain the gap? 40:58 What Chinese job postings reveal about their labs 48:13 Are open or closed models safer? 58:10 How Epoch AI ticks 1:00:55 What a frontier post-training recipe looks like He's one of the people who gives the best feedback on my writing, so I was stoked to have him on.
Show more
Strong agree. There’s way less of this than I would expect — partially due to the fact that making a high quality environment usually involves fairly expensive verification (testing strong models). Still, there can be much more.
Show more
releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now the equivalent of sharing high quality pretraining data but in the new RLVR paradigm
Show more
Real open model nerds knew Xiaomi is cooking with MiMo! An open RL dashboard for the top scoring open model is aura.
I really appreciated this reality check from @natolambert. A lot of AI people underestimate how much the physical world matters, and how little impact AI has had on it so far.
Show more
I was asked by Congressional staff to share my views on open models performance, adoption, and competition vis a vis China. Here’s my briefing with all our latest data and predictions for what comes next. I am very happy to spend time on broader audience content on open models!
Show more
I don’t agree exactly on the assumption that full RSI will work, but this is an excellent video summarizing where we are at.
My best guess is that for scaled RL most of the top Chinese AI labs are starting to use a lot of Huawei for inference and Nvidia for training (maybe not for weird architectures). As agent swarms, even more scaled post-training, etc becomes the norm, this will accelerate their domestic industry.
Show more
Seems like one of the most important research problems for CS academics is llm-supervised peer review. If we don’t solve it the academic institutions are toast. It seems easier than building new institutions.
Show more
Happy to sign this. Truly independent evaluation and oversight of frontier AI is a crucial issue. I think one of the biggest bottlenecks is finding people with the technical capabilities and that high bar of independence. AI Eval Forum is a good group to build initial momentum!
Show more
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: Learn more at
Show more
With the external hack of OpenAI via Claude, closed models continue to be the tip of the iceberg on AI risks, not open models. They have been 1) easier to get started with, 2) more capable & 3) shipped w/ leaky safeguards. Finetuning open models to specific attacks is harder.
Show more
Takeaways from the WSJ article about @HacktronAI using Claude to get into OpenAI's monorepo and issue a pull request (before stopping and claiming their bug bounty) - how many nation states have already broken in and gone much further and stolen a) algorithmic secrets and b) model weights or c) gotten access to user data; these are fair public interest questions - how many have implants in / are dwelling in the openai network as I tweet this - how pervasive is this level of softness to pentesting across all the labs and how far are the labs from the right operating point in the security and r&d friction trade space (probably pretty far it seems) - given the hacktron folks used anthropic's models to pull this off what's the real public safety ROI of anthropic's cyber guardrails; they add friction for legitimate cyber defenders (like our developers at my startup!) but it appears with a bit of work you can use them in actual breaches as happened here - as the article says, the wsj folks and @S1r1u5_ had me do neutral technical review of the kill chain here pre publication; impressive from a human angle (@HacktronAI reminds me of the best of my generation of hackers I looked up to as a kid!) but also from what the models can do; between this and the openai/hf thing, I'm emotionally in a place where I feel my world as a security person is being turned upside down in slow motion with respect to what's coming - elite persistent hacking is becoming rapidly democratized. this is coming like a freight train. we need to harden the world's code and infra as fast as possible and today's non automated methods don't stand a chance of cutting it; the world is a soft target. I continue to be unsettled but very glad I left my comfortable job at Meta to do our automated posture hardening startup
Show more
What it looks like when yelling too loudly about AI risk too early may start to hurt your plans.
This is wild. (I would have said "very small risk.")
At this point, "researchers" are basically DDoSing classic academia. It's pretty sad to see. Either someone comes up with, and executes, a CloudFlare for academia, or it's toast. And the only thing i can think of that might have a chance to scale is more automation combined with a credit/point system like we've discussed a few times here on X over the years, but it needs simultaneous buy-in and coordination from all ~10 major AI conferences and journals, so it's a heavy lift. But if nothing big happens, i think it's game over soon.
Show more
One of the coolest at-scale RL resources made public yet! You love to see it.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
Show more
An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! Called "Never Give Up"
Show more
Join The AI Book Club for a live conversation with @natolambert about his new book, Reinforcement Learning from Human Feedback! 📅 Sep. 24, 9 AM CT 🌐 Online 👇 RSVP
I made a list of the 50+ best articles and resources I’ve read and relied upon in the last few years on open models and open-source AI. Happy reading :)