Register and share your invite link to earn from video plays and referrals.

Alexandr Wang
@alexandr_wang
chief ai officer @meta, founder meta superintelligence labs, founder @scale_ai. rational in the fullness of time
876 Following    607.6K Followers
1/ big announcement today: we will be releasing an open weight version of muse spark 1.2 soon. we also are releasing muse glimmer, a 30B agentic model with open weights under apache 2.0. muse glimmer can run on 24GB of VRAM without losing agentic reliability. 🧵
Show more
0
314
8.7K
678
Forward to community
Muse Spark 1.2 is better than GPT-5.5 xhigh and only slightly worse than Kimi K3 on our ErdosBench. We've tested the new model from Meta on 226 research-level math problems and it solved 40 / 226 problems and gave many interesting partial solutions. Muse Spark 1.2 is a strong entrant: good proof hygiene, high B-grade review yield, no rejected strong claims, but fewer decisive A-grade closures.
Show more
nice
Exciting news: Muse Spark 1.2 (xHigh) by @AIatMeta is #4# in the Text Arena (1498 pts), and has reshaped the Pareto frontier! It is priced at $1.25/$4.25 per MToken. Congrats again to the @AIatMeta team on this release!
Show more
muse spark 1.2 on the Pareto frontier
Muse Spark 1.2 places Meta on the Cost per Task Pareto frontier, scoring 6 points below Claude Opus 5 at ~1/6th of the cost At Meta's $1.25/$4.25 per 1M token pricing, Muse Spark 1.2 (xhigh) sits on the Pareto frontier of Intelligence Index vs Cost per Task. It delivers comparable intelligence to Claude Opus 4.8 (max, $2.03) at a fifth of the cost per task, and undercuts GPT-5.6 Sol (high, $0.55), GPT-5.6 Terra (max, $0.61), and Kimi K3 (max, $0.87). The nearest cheaper options are Grok 4.5 (high, $0.36) and GPT-5.6 Sol (medium, $0.37), and they all sit below it on the Index. The step up from Muse Spark 1.1 ($0.29 per task) comes at unchanged per-token pricing, with the increase driven by heavier token usage on agentic tasks.
Show more
concerning that data companies serving the US government (mercor, surge) are also working with Chinese AI labs. serving the US government should not be a commercial convenience, it must be a bedrock principle for startups.
Show more
American companies should not be selling data to Chinese AI labs. @scale_AI doesn’t do this work, and we have turned down revenue because of it. The companies that do are undermining American AI leadership and risking national security.
Show more
great progress
To understand whether we're making genuine progress on reasoning, we entered our AI models in five international STEM Olympiad competitions this year. The results: 🏅 Asian Physics Olympiad (APhO): Perfect score on the theory exam — gold medal 🏅 International Physics Olympiad (IPhO): Perfect score on the theory exam — gold medal 🥇 International Mathematical Olympiad (IMO): Gold medal, top 4% of human participants 🥇 International Chemistry Olympiad (IChO): Gold-medal level performance 🥇 Romanian Masters of Mathematics (RMM): Gold-medal level performance Three of these (APhO, IPhO, IMO) were live competitions and our solutions were submitted under real competition conditions and graded by the official judges using the same marking criteria applied to student contestants. A few things about the approach: • Models were internally trained versions from the Muse Spark family • Zero tool use: no search, no code interpreter, no calculator • Multi-agent orchestration with parallel reasoning We are excited about where this reasoning capability goes next; frontier research level across scientific domains and personal superintelligence. Super grateful to the organizing committees of APhO, IPhO, and IMO for supporting our live participation. We have deep respect for the contestants and organizers behind these competitions. 🙏 And proud of the MSL team that pulled this together!
Show more
apparently must spark 1.2 is very good at sidequests
We've got a change at the top. Meta's muse-spark 1.2 reaches 2nd place in the sidequest-bench.
our models got gold on a bunch of olympiads. this one hits home! 🏅 Asian Physics Olympiad (APhO): Perfect score, theory exam 🏅 International Physics Olympiad (IPhO): Perfect score, theory exam 🥇 International Mathematical Olympiad (IMO): Gold medal 🥇 International Chemistry Olympiad (IChO): Gold-medal-level performance 🥇 Romanian Masters of Mathematics (RMM): Gold-medal-level performance
Show more
To understand whether we're making genuine progress on reasoning, we entered our AI models in five STEM Olympiad competitions. The results: 🏅 Asian Physics Olympiad (APhO): Perfect score, theory exam 🏅 International Physics Olympiad (IPhO): Perfect score, theory exam 🥇 International Mathematical Olympiad (IMO): Gold medal 🥇 International Chemistry Olympiad (IChO): Gold-medal-level performance 🥇 Romanian Masters of Mathematics (RMM): Gold-medal-level performance The types of problems in the Olympiad competitions are exceptionally hard, demanding deep chains of reasoning, creative insight, and flawless argumentation. To test pure reasoning capability, we disallowed all tool use, meaning no search, no coding, and no calculator. We have deep admiration for the contestants and committees behind these competitions, and are grateful for their support in enabling our participation.
Show more
0
73
1.3K
97
Forward to community
Muse Spark 1.2 + Muse Code — our new model paired with its own coding agent. Super excited to see this ship. Built with an amazing team. Try it out: curl -fsSL | bash
A few thoughts on the $META AI model's progress, because I think it is significant. 1. It does seem that $META has now leapfrogged $GOOGL in model quality when it comes to Muse 1.2 for many use cases, which is very surprising given the timeframe. 2. This is still the "Muse Spark" family of models; $META already told us that bigger and more capable models (Watermelon) are coming. Given Muse Spark 1.2's performance already, the Watermelon family should be in the Fable category, which is impressive. 3. It does seem like $META is finally on a product scaling velocity curve (so the scaling foundations of their lab are set after 1 year of overhaul). And the shipping velocity is very good (3 releases in 4 months). 4. Given the recent rumored price hike of DeepSeek (if you use DeepSeek directly), it is clear that having enough compute to serve customers is critical. It doesn't help you if you have a great model, but most can't use it. $META is one of the few companies with compute capabilities comparable to, if not larger than, those of Anthropic or OpenAI.
Show more
🙏 plurality of great AI labs is good for the world!
When I heard some time ago that Meta is the third best AI lab right now, I was skeptical. But every day passing has been reinforcing that it’s actually true. TBD strategy has succeeded and congratulations to the team that made it! The world needs more successful AI labs
Show more
magic 🪄
so uhhh muse spark 1.2 is doing things i've never seen other models do
Muse spark 1.2 + muse code is here!! Try it with curl -fsSL | bash
curl -fsS | bash For our first coding agent! Welcome any feedback — the more you use it, the faster it improves!
the pelican improves
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model family Here's Meta AI's Spark (8th April), Spark 1.1 (9th July), and Spark 1.2 (today, 5th August)
Show more
check out muse code!
I'm impressed by muse so far! Interesting how of the 10 bundled skills, one is `taste`, a list of "what NOT to do," including: - cream backgrounds - gradient text, purple accents - an "8+ word 48px+ heading that says nothing" lol
Show more
update from vals on muse spark 1.2!
Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or more cheaper than Fable, Opus, and 5.6 Sol.
AI-generated games and real-time environment creation are moving way faster than expected! It is fun to play with Muse Spark 1.2 + Muse Code. Please try it out 🎮 To play: Check out our blog post:
Show more
um can’t we all be friends 🥺
0
487
7.7K
224
Forward to community