Register and share your invite link to earn from video plays and referrals.

Soumith Chintala
@soumithchintala
Building new things @thinkymachines. Also dabble in robotics at NYU. Cofounded @PyTorch. AI is delicious when it is accessible and open-source.
1.3K Following    317.7K Followers
Our plan keeps open weights and safety compatible with each other. We hope that this will help resolve debates and policies that pit open weights and safety as opposite and incompatible; ensuring a future full of open weights!
Show more
Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages.
Show more
Thinking Machines just released Inkling-Small: 276B total, 12B active. A faster Inkling that matches or beats its 975B big brother in many benchmarks. To test its speed, we plugged it into HF's speech-to-speech. Audio goes directly into the model, and it replies using faster-Qwen3TTS. Running on 8x RTX Pro 6000 Blackwell, we get audio back in under 500ms. The normal Inkling needs 2TB of VRAM, this one fits on one node. Because the model hears the audio instead of a transcript, it can hear your tone and emotions. You can check it in the video. Or just go and try it in the space! Really fun model:
Show more
Inkling-small. 2 weeks after inkling Nearly as good as Inkling but 4x smaller. We're just getting started...🔥
Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. Fine-tune it on Tinker today, or chat with it in text, image, and audio on Tinker Playground.
Show more
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
0
16.2K
173.2K
29.8K
Forward to community
it's ironic that the first autonomous AI attack was done by a close weight model defended by an open weight model, where everyone was expecting the opposite
0
116
4K
479
Forward to community
this looks pretty good for agentic. that it fits on a dgx spark is **chef's kiss**
Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark. Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface
Show more
wow, what a world-class model! congrats to the Kimi team.
Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: 🔗 Tech blog:
Show more
0
27
1.8K
49
Forward to community
A model drop by Thinky 🚨🚨 Have been doing some early testing on the model for the past couple of days. Here are some of my findings 1. The reasoning is sharp and concise! Always love to see models that dont ramble 2. Tool calling is beautifulllllll, Its consistent, clean, very well designed for streaming , multi-tool call per step, etc.! 3. Holds up impressively on agentic tasks. In my testing, I was particularly impressed with its ability to run long-horizon tasks with solid error recovery. 4. Most importantly, it was built from the ground up for post-training and customization and this is a real unlock for teams building on open source!! This will be greatly beneficial for continual learning workflows at @trajectorylabs! Amazing step for American OSS models, really excited to keep "tinkering" with it lol
Show more
Modal trained a DFlash speculator that's much faster than MTP, making it a great boost for inference speeds!
Inkling by @thinkymachines is now available on Modal, backed by a custom DFlash speculator for 67% higher throughput and interactivity. Running on Modal Auto Endpoints with SGLang today.
Show more
Inkling is our first open model from @thinkymachines and is now available on Tinker! Check out these quotes from Tinker customers on their experience with Inkling: @_Mantic_AI: "Not only does Inkling outperform Kimi K2.6 on our forecasting evals, it does so with half the output tokens." @trajectorylabs: "We’ve been impressed by how sharp and efficient the model is. Its reasoning is concise, its tool calling is consistently strong, and it holds up well on complex, long-horizon agentic tasks. It feels like a meaningful unlock for what teams can build with open-source models designed for customization." @lightningrodai: "We came away impressed by the model’s underlying reasoning ability. It’s thoughtful, original, and refreshingly unsycophantic.”
Show more
BREAKING: Inkling by @thinkymachines is 9th overall on Agentic Web App Arena by Design Arena with an Elo of 1257 It's an open-weight model in the same performance band as Claude Opus 4.6 by @AnthropicAI and Gemini 3.5 Flash by @GoogleDeepMind This makes Inkling the highest-ranking US-based open-weight model for agentic workloads, achieving frontier-level performance Congrats to the @thinkymachines team for this achievement!
Show more
What do we do at @thinkymachines: Personalization/sovereignty, Human Participation, Decentralization. Democratize AI and make it useful for people. All three of them reduce society's dependence on centralized AGI companies (including ours when we get important), and that is a future worth aiming for. You've seen a preview of this with Tinker, Interaction models and our research openly published on Connectionism. A **lot** more to come very very soon...
Show more
We're building AI that people and organizations can shape and make their own. AI should extend our will and judgment instead of neglecting it; enabling that is the technical challenge we are working to solve.
Show more
Bridgewater, one of the worlds largest hedge funds, a Tinker customer talks through how they've carefully fine-tuned a model focused on what makes interesting financial news. Their fine-tuned model is more effective and cheaper than any frontier model.
Show more
Sorting which financial docs are worth an analyst's time is surprisingly hard for frontier LLMs. With an expert-labeled dataset and on-policy distillation, Bridgewater fine-tuned a model to do it reliably and cheaply.
Show more
sometimes , it really does take a decade for me to understand a paper and appreciate its insight and foresight. <Discovering Causal Signals in Images> is one such paper. wow ... david, @robertnishihara , @soumithchintala , @bschoelkopf & @LeonBottou really did see the future.
Show more
Cluster magicians and GPU whisperers, come join us! We’re looking for supercomputing engineers to build the infrastructure behind real-time interactive models, Tinker, and large-scale training: scheduling, storage, networking, reliability, and distributed systems at scale. Hiring in NYC and SF
Show more
thinking machines....the people are incredible
0
144
3.3K
72
Forward to community