Register and share your invite link to earn from video plays and referrals.

Mira Murati
@miramurati
Now building @thinkymachines. Previously CTO @OpenAI
644 Following    1M Followers
We’re building machine intelligence to expand human will & judgement. Glad to have you on the team to help us figure out how to make this future safe.
I’ve joined @thinkymachines to work on safety & alignment. The default trajectory is that as AI gets more powerful, control over it will concentrate in the hands of a few. I’d rather build safety systems that let control over AI be shared widely. Longer thoughts below.
Show more
An important conversation about where AI is going and the work that is cut out for us.
Our own @johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.
Show more
0
48
1.5K
96
Forward to community
Pleasant surprise to see this. Paul is visionary and principled, and I've admired his work for over a decade
I think it's possible to make a lot of progress toward models that are both open and safe! The former will become more important to avoid centralization & spread benefits, and the second will become more important as *gestures at state of things*
Show more
We built Inkling to be fine-tuned, and @besimple_ai did an impressive job of that. Just 25-100 hours of specialized audio takes Inkling to the top rank for recovering subtle details from speech.
We published how @thinkymachines 's Inkling did on Voice Code Bench few weeks ago, and then we asked, what if we fine-tune Inkling? Inkling's main value prop is its small size and robustness, making it a good base model for domain specific fine-tuning. So we that's what we did. We fine-tuned the same Inkling speech model on 1, 25, and 100 hours of our proprietary data containing alphanumeric entities. And the curve kept moving. 🍊 Not surprising, on the standard 300-item VoiceCodeBench evaluation, the 100-hour checkpoint delivered the strongest result. Untrained Inkling → 100-hour post-trained Inkling: 📈 Task Success Rate: 56.33% → 79.00% (+22.67 pts) 📈 Entity recovery (CTEM): 86.84% → 94.80% (+7.96 pts) 📉 VoiceCodeBench WER: 2.3748% → 1.6107% (32.2% relative reduction) 🔧 139 misses fixed, 21 prior hits regressed: net +118 exact entities recovered And the gains scaled as data scaled: 1h: 88.06% CTEM / 59.33% TSR / 2.8188% VCB WER 25h: 92.85% / 72.67% / 1.8517% 100h: 94.80% / 79.00% / 1.6107% The 100-hour model recovered values the base model missed: •--revert-last, instead of splitting one flag into “--revert --last” • tests/auth/login.spec.ts, instead of test/auth/login.spec.ts • ALLOWLIST_CIDR, instead of inserting an extra underscore • SN-7KX-9042, instead of dropping the final digit The largest entity-type gains were email addresses (+30.8 pts), postal addresses (+30.0), file paths (+23.5), environment variables (+22.9), and IP addresses (+20.0). That’s the Besimple thesis: targeted human data can move the production metrics that matter for voice agents, even when the base model is already strong. DM me if you want to evaluate your model on this benchmark or build the data that moves it. 🍊 #SpeechRecognition# #VoiceAI# #ASR# #Transcription# #DataQuality# #PostTraining# #Benchmarks# Checkout the full blog at:
Show more
We’re growing our evaluation team at @thinkymachines. We care a lot about whether models are actually useful and whether our measurements are good enough to serve as the y-axis for scaling research. We’ll work across a wide range of problems and support several workstreams, including building signal-bearing internal evals for predictive scaling, turning real user feedbacks and product workflows into evals that close usability gaps, auditing graders and harnesses, and developing new benchmarks for customizability. If you want to work on this, apply here:
Show more
Cleaning data and aligning the reward function for RLVR takes expertise and effort upfront, but the result is a model that's state-of-the-art on a complex task. Guest post by researchers at UIUC and Bridgewater, in collaboration with our team.
Show more
LLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker, @maxYuxuanZhu and @ddkang (UIUC and Bridgwater) trained the first text-to-SQL model to beat the human mark.
Show more
Quick, accurate summaries make an even better resource and more easily searchable by both people and LLMs. We're glad to see Inkling contributing!
We use Inkling-Small to turn paper abstracts into quick, useful summaries. open weights × open science 🤝
Are you doing AI safety research that could be accelerated by up to $50k in @tinkerapi credits? If so, we want to hear from you. 🦺
We want to improve Inkling’s agentic performance. To help us understand its real-world behavior, we are making it available for free on OpenRouter (only with agentic harnesses) for the next few weeks, starting now. We’ll use the data, disassociated from accounts, to better it.
Show more
0
56
2.1K
126
Forward to community
Alex is good to work with, can kick a ball and captain a football team pretty bloody well, and has often-good thoughts on safety. If you want to join an ambitious and pragmatic AI safety team, you should get in touch :)
Show more
We share our approach to open-weights releases, how we assessed Inkling, why safety depends on both the model and the ecosystem it enters, and how testing, staged access, and stronger defenses can create a path toward greater openness.
Show more
Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages.
Show more
0
52
1.6K
116
Forward to community
Whereas I felt like it took a village to release inkling, inkling-small felt much more routine 😆 We just took the pipeline used for Inkling, passed in a smaller model, and voila - new model! Inkling small benefited quite a bit vs Inkling from some minor improvements, but there's still so much more left in the tank...
Show more
Thanks @ArtificialAnlys for working together on evaluations! Excited to see Inkling-small on the pareto line🫡
Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward to seeing what people make with it.
Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. Fine-tune it on Tinker today, or chat with it in text, image, and audio on Tinker Playground.
Show more
0
142
3.7K
281
Forward to community
The knowledge that makes AI useful is diffused. It lives with scientists, engineers, clinicians, firms. For AI to benefit from distributed knowledge, it must itself be distributed. Agree with Jensen that this is a future worth building.
Show more
0
328
8.9K
1.1K
Forward to community
It truly takes a village to release a model, perhaps especially an open weights model. Actually doing the entire process from scratch, from data to pretraining to posttraining to actual release, gives a lot of appreciation for anyone who does it! There's so many places to go wrong, and indeed so many things we would (and will!) do differently for a new model. But I'm happy with where we ended up :) Onto what's next!
Show more
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Show more
0
50
1.3K
45
Forward to community
Our first model, Inkling. Trained from scratch, weights are open, fine-tunable on Tinker today.
0
399
13.7K
1.1K
Forward to community
Today we share the worldview behind our mission. Human values don't average out. Local knowledge can't be centralized. The good future has many AIs, raised in different places, shaped by the people they serve, disagreeing with each other the way we do.
Show more
0
149
4.5K
526
Forward to community
The video of my Stanford CS25 guest lecture, From Language Models to Native Multimodal Intelligence, is now online. I discussed how the core ideas behind LLMs has shaped multimodal AI, from architectures to training paradigms and scaling, and where the next challenges may lie. 🧠🌐 🎥:
Show more
0
23
1.7K
197
Forward to community