Register and share your invite link to earn from video plays and referrals.

Nathan Lambert
@natolambert
Open model research @ something new. Prev. co-led Olmo at Ai2. Contact via email. Writes @interconnectsai Wrote The RLHF Book, 🏔️🏃‍♂️
931 Following    96.8K Followers
It's a good time to share a new @interconnectsai project we made to help make sense of the accelerating open-weight releases these days. The Artifacts Hub builds on our monthly open model roundups and daily monitoring of every model on @huggingface. The new free resources are: 1. The Artifacts Hub — a curated view of the models trending on Hugging Face, highlighting inference tokens via Open Router, model intelligence via Artificial Analysis, and our tailored adoption metrics building on top of Hugging Face’s data. 2. Our Adoption Dashboard — a living dashboard of download and derivative model numbers by geography and organization. This highlights the US-China gap and growing players in the open ecosystem. To date, our primary efforts on Interconnects have been release recaps for popular models like Kimi K3, GLM 5.2, DeepSeek R1, etc. and monthly round-ups of the open models that matter, Artifacts Log. We’re expanding on these, building on the tools and internal data we’ve collected for other projects like The ATOM Project (and report). This allows us to capture our ecosystem view of open models, develop methods for understanding adoption of giant MoE models, and everything in between. We’re sharing them freely to help the open ecosystem find its strengths and grow. The Artifacts Hub right now covers 792 models released in the last two years, across the core text-focused language models and multimodal generative models. At Interconnects we follow the data of every model on Hugging Face, analyze the core few thousand LLMs (this list is public on GitHub and regularly updated), and hand select these core few hundred for further explanation. For the most popular models, the Hub let’s you quickly see how far behind the model was in terms of frontier intelligence based on Artificial Analysis’s Intelligence Index, compare Hugging Face and Open Router adoption to similar models, glance at relative adoption metric (RAM) scores for time-size normalized downloads, or look at the VAIL similarity index of models with related generations. A snapshot for what you’d see for something like GLM-5.2 is below. Thanks to @mnshah at VAIL for encouraging us to make this and @huggingface, @OpenRouter, & @ArtificialAnlys to making such useful data openly available.
Show more
Kimi K3 license. It's inspired by MIT but distinctly non-commercial, where any company making over $20M/yr must get a specific commercial deal (and display Kimi K3 if over 100M users or $20M/mo revenue)
Show more
0
205
3.9K
276
Forward to community
To anyone that accuses me of being a China shill on distillation -- I've been worried about this for a long time and we're just starting to feel the pain of being behind on open models, and it'll only get worse from here if we don't find ways to support building them in the US.
Show more
My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024. Transferring as much of the intuitions of building Olmo as I possibly can in the book format. The book is launching with an over 10 hour, full course with slidedecks, functional code for the training chapters, an example model completions library, and of course the free online web version. Physical orders from Manning will ship in 1-2 weeks, and Amazon a week or so after. Thanks for your support!
Show more
0
144
3.2K
267
Forward to community
Important to read. China is committed to continuing its open, global ai approach.
jUsT gOoD bEcAuSe Of DiStIlLaTiOn
0
75
1.6K
86
Forward to community
Claude Fable is another big step in being able to make nice lectures based on existing educational content. Much better than Opus. GPT 5.6 is still very far off here. Is a good example of where Claude Code being a bit easier to work across different knowledge work tasks.
Show more
Another insane own goal on AI research leadership from the U.S. It's like "how do I remove myself from an international scientific network" 101 class
Utterly self-defeating. “The National Science Foundation (NSF) has decided to ban collaborations between every U.S. scientist it funds and nearly all Chinese research institutions and their employees.”
Show more
The open model community is extremely unprepared for when a model gets stuck in the undefined white house licensing regime - and it could permanently knee cap the open model economy within 6 months. Why this'll happen and what we can do:
Show more
I got around to reading ai 2040, and I’m very happy with their focus on transparency (which I strongly agree with), but I find it odd to talk about that so much with no real discussion of open source.
Show more
Let’s goooo USA 🇺🇸🦅🏆
Trying to help a close collaborator get their @X account back after falling for the AGI-level phishing campaings, but they're not famous like me, can someone help do me a favor? Will check DMs :)
Selling tokens and building the token machines are both going to be great businesses. I'm confident in more companies that build open models figuring out how to make money. Open models bring attention and attention is a funnel to revenue for many companies today.
Show more
Moonshot AI's Kimi has reportedly hit $300 million ARR as of mid-June, with API revenue exceeding 70% of total. A new funding round is underway at $31.5 billion pre-money, per Chinese financial media. Four months ago, the valuation was $10 billion.
Show more
When we were in China, @xeophon and I made a quick detour to visit Meituan. They continue to be one of our favorite open model builders, as they're showing how a variety of companies can succeed here and baffle a lot of people as to why they're making models. Meituan is one of the larger tech companies in China. They're building LLMs to add services to their own products. In China the notion of the "super app" is very popular, so this dream of more services for users with AI is very natural there. With this, Meituan wants to own the full stack of how they deliver value to their users. When we visited, they were very unassuming about everything. We just met a few people from the LLM team, a quick meeting about building models. They build general foundational reasoning models, and then fine-tune it further for their products. They can release the general model to support the ecosystem and learn how it can be used. Their focus was very clearly on ownership, and a hint of cost-saving, so the recent news of v2 being trained on asics fits with that mentality. They want to deliver real products to users with low cost. Companies like this will keep building models in China. It's a small micro study of how different the players in the AI ecosystem are. Kimi, Z ai, etc are all much flashier offices, come across as the "hot new thing" but Meituan has the talent and resources to build models as well. Congrats to the Meituan team & thx for having us!
Show more
Introducing LongCat-2.0 🐱 1.6T parameters · MoE with ~48B active · 1M context The full model behind Owl Alpha on @OpenRouter — now available. Built for agentic coding from the ground up: ◆ LongCat Sparse Attention (LSA) — scales efficiently for 1M-context tokens ◆ Zero-Compute Experts — dynamic activation 33B–56B per token, zero wasted compute ◆ MOPD — three specialized expert groups (Agent / Reasoning / Interaction), gate-routed per task How it stacks up: → Terminal-Bench 2.1: 70.8 → SWE-bench Pro: 59.5 (GPT-5.5: 58.6) → SWE-bench Multilingual: 77.3 → FORTE: 73.2 · RWSearch: 78.8 · BrowseComp: 79.9 📖 Tech Blog: Try it across different scenarios 🧵👇
Show more
letssss gooooo breaking this bad boy out today loooooooooooong cat
Introducing LongCat-2.0 🐱 1.6T parameters · MoE with ~48B active · 1M context The full model behind Owl Alpha on @OpenRouter — now available. Built for agentic coding from the ground up: ◆ LongCat Sparse Attention (LSA) — scales efficiently for 1M-context tokens ◆ Zero-Compute Experts — dynamic activation 33B–56B per token, zero wasted compute ◆ MOPD — three specialized expert groups (Agent / Reasoning / Interaction), gate-routed per task How it stacks up: → Terminal-Bench 2.1: 70.8 → SWE-bench Pro: 59.5 (GPT-5.5: 58.6) → SWE-bench Multilingual: 77.3 → FORTE: 73.2 · RWSearch: 78.8 · BrowseComp: 79.9 📖 Tech Blog: Try it across different scenarios 🧵👇
Show more
I feel like the Chinese labs I visited felt like this. It’s what Ai2 felt like. It’s the most reliable energy for making something amazing (even when you’re an underdog with the number of resources). I’m always super excited to find my next star intern.
Show more
Founder Tip: Load up on intern energy and naivety. Today's interns are a question away from most knowledge thanks to LLMs, and they haven't yet learned that what you're asking them for is supposed to be impossible.
Show more
With everything going on, it gives me hope that there's such a diversity of companies building open models today. A lot of the story of open models unfolds under the shadow of the biggest frontier models. Lots of unearthed value.
Show more
Artifacts 22: Zyphra, Cohere, Poolside, and others are expanding the breadth and diversity of the ecosystem. In this issue, a total of 30 models from may/june you should be aware of, from: NVIDIA @NVIDIAAI (3) Cohere @Cohere_Labs (2) Zhipu @Zai_org Zyphra @ZyphraAI (3) Poolside @poolsideai Moonshot AI @Kimi_Moonshot StepFun @StepFun_ai Dolphin @dphnAI Google @GoogleAI (3) Nex AGI @NexEcosystem Liquid AI @liquidai MiniMax @MiniMax_AI Swiss AI Initiative @apertusllm JetBrains @jetbrains Microsoft @Microsoft H Company @hcompany_ai Datalab @datalabto (2) Baidu @Baidu_Inc PaddlePaddle @PaddlePaddle Ideogram @ideogram_ai KREA @krea_ai Photoroom @photoroom_ML Read the issue below.
Show more
This is real and a horrible consequence of vibe regulation of frontier models.
Getting regulated by a government because your model is "too dangerous" is the best marketing (especially for enterprise sales) so everyone is trying to get it now 😅😅😅