As a test of our progress to advance the frontier of AI research, in June we entered the next generation of our autonomous AI research system, AIRA₃, in a live Kaggle competition run by NVIDIA to fine-tune a 30B Nemotron model. The challenge was to teach the model to reason better — all competitors had access to the same information and were graded externally on a private test set.
AIRA₃ placed 8th out of ~4,000 teams to win Gold, outperforming human competitors who had access to the same frontier tools.
We believe this is a reliable signal that AIRA₃ can improve a targeted capability of an AI model at a level similar to human experts.
Show more
We’re excited to release Muse Spark 1.3 with improved performance on agentic and coding tasks, and a focus on real-world usability.
Key capabilities:
→ Sustains longer-horizon work across multiple workflows in a single thread
→ More actively collaborates with users: it asks clarifying questions, flags when it's stuck, confirms before consequential actions
→ Better calibrated on its own limits instead of hallucinating outcomes
→ ~20% fewer tool calls and ~25% fewer tokens vs. Muse Spark 1.2 in internal comparisons
Show more
Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API.
Next up 🍉 and Muse Spark open weights releases coming soon.
Show more
Today we're releasing Muse Spark 1.3, our strongest model for agentic and coding tasks in the spark model line. It’s designed to support longer-horizon work, excel in agentic tasks, and follow complex instructions more reliably than previous models. Muse Spark 1.3 is available today in Muse Code and Meta Model API.
Show more
Introducing Muse Voice Transcribe, the first real-time audio perception model from Meta Superintelligence Labs.
Muse Voice Transcribe delivers real-time streaming ASR, diarization with 20+ speakers, and endpointing. It’s multilingual with seamless code-switching and improves accuracy with language, keyword, and context biasing.
The model ranks first on
@ArtificialAnlys streaming speech-to-text and on public diarization benchmarks.
Show more
Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use.
Today, we’re sharing new evals and demos that illustrate the breadth of the model’s visual understanding and reasoning capabilities.
Let’s start with a demo that shows how Muse Spark parses multimodal observations and calls tools to guide a robot to navigate in an unstructured environment to find a rubber duck.
🧵👇
Show more
1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.
Today we're releasing Muse Code (beta) and Muse Spark 1.2. Muse Code is our first coding agent — it coordinates multiple persistent subagents to complete complex engineering tasks faster, more accurately, and with less intervention. Muse Spark 1.2 is our newest model, co-trained with Muse Code for best performance when paired together. Available now through Meta Model API.
Try it: curl -fsS | bash
Show more
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
Show more
Great writeup explaining our stance on the future after superintelligence -- empowerment and free will for everyone rather than a future designed by a handful of institutions.
I wrote about why we believe the future is for everyone. More coming about a positive vision for a world with superintelligence soon.
An internal muse spark achieved a perfect score on APhO. Big congrats to the team, and thank you to the APhO committee!
To demonstrate Meta AI's advanced reasoning and multimodal capabilities, we submitted a model to participate in the Asian Physics Olympiad’s theoretical exam. We’re happy to share that our model achieved a perfect score of 30/30, tying with the top 3 student contestants.
We appreciate the APhO committee for letting our model participate in the competition:
Show more
Excited to share what we’ve been building at Meta Superintelligence Labs!
Today we’re launching Muse Spark 1.1, our strongest model yet for complex agentic workflows — delivering massive gains in agents, computer use, coding, multimodal reasoning, and multi-agent orchestration.
This is also our first API release! We’d love to hear your feedback.
And we’re just getting started, super proud of the team behind it, and the larger models are training right now. 🚀
Show more
Today we are launching Muse Spark 1.1, an upgrade to muse spark 1 that greatly improves agentic, coding, multimodal, and computer use capabilities. We're also launching the Meta Model API in public preview.
Show more
1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8.
available now through the new meta model api and in meta ai. 🧵
Show more
(1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
We are excited to launch muse Image, the first media generation model built by MSL. Muse image can use reasoning, refinement, and tools to improve precision and quality, with very clear test time scaling trends. Also sharing Muse Video preview today.
Show more
Today we’re introducing Meta AI Voice Conversations powered by Muse Spark that let you talk naturally to Meta AI (interrupt, switch topics, or swap languages), and as you talk, Meta AI can generate images and pull up recommendations from Reels, maps, and more. We’re also bringing live AI to the app, so you can point your camera at the world and ask about what you’re seeing in real time.
Show more
this is a very deep write-up:
a good writeup about Muse Spark on a few complex queries (multimodal, stock analysis, coding):
We are committed to responsible deployment of the technology and have a really good safety team to back it
1/ Muse Spark is live, and alongside it, our new Advanced AI Scaling Framework which details how we evaluate and prepare for advanced AI. We tested across bio, chem, cyber, and loss of control risks before and after mitigations. Muse Spark achieves a 98% bioweapons refusal rate on BioTier-refuse, highest across the models we benchmarked.
Show more
Humble beginnings lol
when msl first started we had 4 researchers, 1 gpu, and
@jiahuiyu was wearing a facemask all the time from going viral. we’re so back!!