Congratulations Levent, Tristan, and OpenAI! What a miraculous time to be alive!
I view this as the first successful achievement of Recursive Self Improvement or RSI. Models get increasingly better at Math and get increasingly used by mathematicians to solve all kinds of difficult problems in their domain, thereby contributing very rare and very hard tokens that are then used in the model training to further improve the model's math capabilities: generalizing and drawing connections across sessions and everything else the models learn from; which is then served back, and so on and so forth.
This happened first for math but will eventually happen for every domain.
Show more
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work.
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
Show more
Surprised to learn today how many people did not know that if you don't turn off "Improve the model for everyone" in chat then your data will very likely be trained on. Also, its on by default. Same applies to codex.
Show more
Fun fact: RLSlow was named after Thinking, Fast and Slow. The idea was that language models already had a kind of "fast" thinking, producing an answer immediately, and that we could use RL to teach them "slow" thinking: deliberate, multi-token reasoning that spends more compute working through a problem.
What I remember most from those days is how early the team developed real conviction in the direction, and how much work went into earning it. We were developing the algorithms, designing careful experiments to test the ideas, and watching the empirical evidence accumulate. These were many long nights babysitting runs, understanding what the results were telling us, and figuring out what to try next. A lot of those early discussions were with Ilya, and later with Jakub.
Pretty early, the evidence had already pushed us to a strong view: RL for reasoning would scale. Models would learn to spend more compute at inference time to reason through increasingly hard problems, and this would fundamentally change how we think about inference.
Many of us also spent countless hours reading through reasoning traces. Seeing how models arrived at answers, not just the answers themselves, felt like a powerful new lens on generalization and alignment.
These ideas feel obvious now. They really weren’t then.
Show more
Excited to see Muse Spark 1.3 out in the world. There’s a lot of progress in this release, and with 1.3 Max specifically, the intelligence you get per dollar is kind of ridiculous!
If you’re a developer and haven’t tried Muse Spark yet, you’re running out of reasons not to.
Show more
Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API.
Next up 🍉 and Muse Spark open weights releases coming soon.
Show more
Since
@shuchaobi is posting pictures here’s one.
Fun fact — we literally went to the beach when the model was solving IChO. We trusted it enough. This pic was taken then. Over the years it’s stuff like this that one remembers :)
I wonder if other labs got their Golds with the same panache 😇
(with
@TrapitBansal)
Show more
Students who earn gold at these Olympiads are exceptional and often go on to make major contributions in their fields. As we continue to scale these models, it’s exciting to imagine what one that excels across physics, chemistry, and mathematics could achieve.
3/5
Show more
We have been investing in advancing multimodal reasoning for Muse Spark and we participated in this year's major science olympiads. It was great to see our model obtain a 100th percentile score at the International Physics Olympiad (🥇), International Chemistry Olympiad (🥇) and the Asian Physics Olympiad (🥇) in theory exams!
Show more
It’s been a very fun making this happen at TBD, when the lab itself is just about a year old.
Real pleasure to work
@TrapitBansal @shuchaobi and many other friends and collaborators on this!
Students who earn gold at these Olympiads are exceptional and often go on to make major contributions in their fields. As we continue to scale these models, it’s exciting to imagine what one that excels across physics, chemistry, and mathematics could achieve.
3/5
Show more
Every parent I know here has the same problem: finding something to do with their kids and dog that still feels new, while threading limited weekend hours, drive time, weather, and nap schedules. It's basically a second job!
I wanted my sacred family time spent making memories, not scrolling for plans. So I built : one plain sentence in, a whole day out. kids, dogs, naps, all of it.
It's live, go plan something 🎠
Show more
Introducing Phirky: your bay area weekend planner.
Type it like you'd say it: "rainy Saturday, toddler under 3, under $20, dog-friendly, near Palo Alto"; and hands you a ranked plan. Every place human-verified: realistic hours, parking, shade, water bowls.
Show more
Asian Physics Olympiad 2026 was a good eval of our model's multimodal and reasoning capabilities as we continue to scale. Happy to share that the model achieved a perfect score!
To demonstrate Meta AI's advanced reasoning and multimodal capabilities, we submitted a model to participate in the Asian Physics Olympiad’s theoretical exam. We’re happy to share that our model achieved a perfect score of 30/30, tying with the top 3 student contestants.
We appreciate the APhO committee for letting our model participate in the competition:
Show more
I believe that exploring and making mistakes is key to learning and research.