Register and share your invite link to earn from video plays and referrals.

Interconnects AI
@interconnectsai
What you need to know about the latest models and AI research trends, from @natolambert
2 Following    8.9K Followers
Artifacts 23: Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier. In this issue, a total of 24 models from june/july you should be aware of, from: Thinking Machines @thinkymachines (2) Tencent @TencentGlobal (2) Poolside @poolsideai (2) DeepSeek @deepseek_ai Moonshot AI @Kimi_Moonshot Meituan LongCat @Meituan_LongCat Motif Technologies Swiss AI Initiative @apertusllm AMD @AMD Upstage @upstageai InclusionAI @TheInclusionAI Moondream @moondreamai Baseten @baseten IBM @IBM Mistral AI @mistralai Google @GoogleAI Nanbeige @nanbeige Kwaipilot @KwaiAICoder InternLM @intern_lm Microsoft @Microsoft fal @fal Read the issue below.
Show more
Artifacts 22: Zyphra, Cohere, Poolside, and others are expanding the breadth and diversity of the ecosystem. In this issue, a total of 30 models from may/june you should be aware of, from: NVIDIA @NVIDIAAI (3) Cohere @Cohere_Labs (2) Zhipu @Zai_org Zyphra @ZyphraAI (3) Poolside @poolsideai Moonshot AI @Kimi_Moonshot StepFun @StepFun_ai Dolphin @dphnAI Google @GoogleAI (3) Nex AGI @NexEcosystem Liquid AI @liquidai MiniMax @MiniMax_AI Swiss AI Initiative @apertusllm JetBrains @jetbrains Microsoft @Microsoft H Company @hcompany_ai Datalab @datalabto (2) Baidu @Baidu_Inc PaddlePaddle @PaddlePaddle Ideogram @ideogram_ai KREA @krea_ai Photoroom @photoroom_ML Read the issue below.
Show more
Latest open artifacts (#21#): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 & others. On CAISI's V4 assessment. An eventful month with one flagship release after another
Show more
How open model ecosystems compound Further reflections on China's high-participation, open-first AI ecosystem.
Notes from inside China's AI labs Lessons from my trip to talk to most of the leading AI labs in China.
The distillation panic ‘Distillation attacks’ is a horrible term for what is happening right now.
Reading today's open-closed performance gap The complex factors that determine the single evaluation number so many focus on. Plus, how this changes in the future.
New video! Talking through my 10+ open model pieces from early 2026 and how they fit together. They're all trying to figure out where open models go next.
My bets on open models, mid-2026 What I expect to come next and why, focused on the open-closed gap.
One of my key strategies with Interconnects is to develop the practice of making my work obviously compelling to a wider audience, keeping them hooked over time and wondering what I'm up to, etc. It pays off so much on the downstream influence of your work to have direct audiences for it. These types of articles are a pure grind, but the directly convert hard earned reputation into success for projects. Many people, especially academics, too far too many projects and invest too little in each. The amount of work just on distributing each project is fairly high. If you don't work on distributing it, don't expect to get new readers. It'll seem like a lot of your audience already knows everything if you've mentioned it once, but repetition is key. People are busy, algorithmic feeds are fickle, and people don't remember points/brands until they've seen it many times. Do the work (and read my post contextualizing why all my recent efforts are important)!
Show more
What I’ve been building: ATOM Report, post-training course, finishing my book, and ongoing research
Excited to launch the accompanying free RLHF Course for my book. To kick it off, I've released: - Welcome video - Lecture 1: Overview of RLHF & Post-training - Lecture 2: IFT, Reward Models, Rejection Sampling - Lecture 3: RL Math - Lecture 4: RL Implementation I'm going to add question & answer videos throughout the lecture to go deeper on topics that need it, and potentially cover some topics that are too recent and in flux to go in print. I expect 10-15 videos in total over the next few months. At the same time, development around the code for the book is picking up. It's a great time to build the foundation for post-training methods. YT playlist and course landing page below.
Show more
0
50
1.8K
237
Forward to community
The inevitable need for an open model consortium And yes, I hate consortia too.
Claude Mythos and misguided open-weight fearmongering Another dance around fears of open-source.
im once again asking you to think before you ask someone to ban open models
1. dont fall for anti open model fearmongering, but 2. acknowledge that AI capabilities are proceeding fast, and eventually there may be a reason to be more careful with open weight models I don't think Mythos is that trigger, but I'm not 100% confident
Show more
My book, Reinforcement Learning from Human Feedback, is wrapping up and going into final production (copyediting, making pretty, formatting, etc.). Shipping to you in 1-2 months! It's a wonderful project to create a foundation of knowledge for the research communities that I love and operate in. It’s the book I wish I had when starting on my LLM journey about 3 years ago. The book’s deepest cut is on core reinforcement learning methods, intuitons, and implementations for LLMs. These don’t live in isolation, and it’s presented in the broader context of post-training methods and unsolved problems in RLHF. A nice balance of depth and breadth. I’m always asked about the title, and I am staying firm that this is THE book documenting the organization of the field of RLHF. Any other topic is too dynamic, where writing a book today would be immediately outdated. RLHF is largely being overshadowed by lots of other developments in AI, but will always be around and at the forefront of human-AI interactions. The topic deserves coverage in depth and this platform. Thank you for all your support. More projects related to the book being announced soon 🎥 I'm excited to reconnect with the community through in-person book events this summer and fall.
Show more
Check out the latest open model ecosystem data! Excited to build more tools in this space soon.
New report with @xeophon is out with the latest open model adoption data we have gathered for Interconnects & The ATOM Project. At the surface level, we can see Chinese models continuing to accelerate in adoption. The report details much more. 1. We manually curate ~1.5K of the most important language models, creating a specific set of models to focus our analysis on (excludes embedding models, local inference models like MLX/GGUF, etc to have accurate download rankings). 2. Studying other adoption metrics, such as derivative models and inference share on OpenRouter, to show how they correlate with downloads, while often sifted in time. China has a strong lead here too. 3. Better classification of downloads across model sizes. Large models still are the models where Qwen is least competitive, relative to other model builders. 4. Expansion of our Relative Adoption Metric (RAM) to show standout recent models (we'll check Gemma 4 on Friday); Qwen 3.5, Nemontron 3, Kimi K2.5, all showing very strong adoption. Overall, this is another step towards formalizing and making public better data on the open language model ecosystem, so the community can better understand the impact and trends of its adoption. More on this soon!
Show more
Gemma 4 looks great on paper, just what many people need, but mostly I'm focused on the questions around what makes an open model succeed in the long term. We desperately need to build a field of research on "what makes a finetunable model" if we want the open model economy to succeed. Too much of that research happens behind closed doors of companies. Plus, 240 days after The ATOM Project was released, I share reflections on the recovery of American models, and what that means for the next phase.
Show more
Gemma 4 and what makes an open model succeed Hint: it's not benchmark scores.
One of Interconnects' truest missions is improving model licenses and other tiny cuts that get in the way of model adoption. Thank you Google for listening.
Google dropped 4 different Gemma open-weight models! I'm most excited that they're finally adopting a standard Apache 2.0 open source license. This'll massively boost adoption. The standard of better licenses was set by mostly Chinese open model labs, and now labs in the U.S. companies are following suit. The models are really like 31B dense, 26B-4B active MoE, 8B, 5B dense (called smaller for some reason). Base models too. Good sizes for tinkering, some local uses, and research (8/5B). 30B is particularly a great size range for building useful tools (which is why we made Olmo 3 that size too). Gemini doesn't release bad models so I'm excited to try these! Congrats Googlers.
Show more