Register and share your invite link to earn from video plays and referrals.

Shuchao Bi
@shuchaobi
Research at Meta Superintelligence Labs, RL/post-training/agents; Previously Research at OpenAI on multimodal and RL; Opinions are my own.
579 Following    14.5K Followers
SWE-Together Update: Claude Fable 5 and 5.1 are the new top models on SWE-Together, and Meta's Muse Spark 1.3 is the best value on the board. SWE-Together is our benchmark of 109 real coding tasks with a simulated user in the loop, built to tell you whether a new model will actually match your expectations before you hand it your work. Six quick findings from the updated board: 1. Fable 5 beats 5.1 on consistency, but 5.1 is more efficient. At their best the two actually perform the same. The gap is in the worst runs, where Fable 5.1 scored zero on 14 trials and Fable 5 on only 6. However, 5.1 is 25% faster, about 2x cheaper, and needs slightly fewer corrections from the user (1.45 vs 1.53 corrective messages per task) than Fable 5. 2. Muse Spark 1.3 is the value outlier. Per solved task, it is about 2.5x cheaper than Fable 5.1 and 5x cheaper than Fable 5. Performance-wise, it ties with Fable 5 and 5.1 for best on Python (73%) and Rust (100%). 3. Newer is not automatically better at collaborative coding. GPT-5.6-Sol is the newer model, yet it needs more corrections from the user than GPT-5.5 (1.66 vs 1.59 per task). 4. "Stronger models need less steering" is a trend, but not guaranteed. With more models added, the correlation between pass@1 and user correction weakens from -0.92 to -0.72, and Fable 5 takes more corrections than Opus 4.8 (1.53 vs 1.38) despite a 6 pp higher pass@1. 5. Frontier progress contributes greatly to stability. Fable 5 converts 89% of its pass@1 into pass² (both runs solve), versus 60 to 67% for DeepSeek V4 Pro, GLM-5.1, and MiniMax 2.7. 6. There is still plenty of headroom. 16 of 109 tasks have zero solves across all tested models, and even Fable 5 at 69% pass@1 sits about 9 pp below the ~78% that the original human patches scored. The full leaderboard, with per-task and per-trial breakdowns, is at and the benchmark is open source at
Show more
literally making you money, download today and give it a try
$META just went from 3.5% to 45.4% token share on OpenCode in just over two weeks Muse Spark 1.3 being good + free is enough to become the default for most users Default gets you usage → usage gets you data → data makes the next model better Anthropic and OpenAI can’t afford to play the game like this but $META can
Show more
BREAKING: Muse Spark 1.3 (xhigh) takes 1st overall on Website Arena with an Elo of 1362! This is a jump of 5 positions from Muse Spark 1.2, establishing a new Pareto frontier for Speed and Price. Only a month after the release of Muse Spark 1.2, @AIatMeta has topped this category on Design Arena. Note: GPT-6 Astra is still pending final results. Congrats to the @Meta team!
Show more
We've run a bug bounty program on Muse since early in development. Today it becomes public, with published payout guidelines. We pay by the impact demonstrated. More details at
Show more
0
19
1.1K
46
Forward to community
Meet Muse, your personal AI agent. Muse doesn’t just answer questions, it actually does the work across the apps you already use. It helps you stay on top of things, takes tasks off your plate, and turns long-term goals into action plans.
Show more
Another independent evaluation shows just how performant and cost-efficient Muse Spark 1.3 is. The max mode costs much less than GPT 5.6 Sol while achieving similar or even slightly better performance on your real daily work.
Show more
muse is way hotter than the other personal agents (and also way faster, but mostly hotter)
Muse Spark 1.3 Max by @AIatMeta has reshaped the Pareto frontier for Code Arena: WebDev! Meta's latest model at Max reasoning is doing something interesting on the Arena Pareto frontier: it's the only model holding down the wide price band between Qwen3.8-max ($5/MToken) and Qwen3.8-Flash-Next ($0.39/MToken). The next closest model to the frontier is Meta's own, Muse Spark 1.3 on xHigh reasoning. Muse Spark 1.3 Max performs just 20 pts lower than Qwen3.8 (Max) while costing 30% less, and 24 pts from Kimi K3 (Max) while costing 70% less. - Muse Spark 1.3 Max: 1650 pts | $3.50/M - Qwen3.8 (Max): 1670 pts | $5/M - Kimi K3 (Max): 1674 pts | $12/M Overall in Code Arena: WebDev, Muse Spark 1.3 Max landed #8#, ahead of Claude Fable 5 at #10# (+22 pts), Grok-4.6 (High) at #12# (+26 pts), and GPT-5.6 Sol at #14# (+33 pts). Congrats to the @AIatMeta team on this release!
Show more
try it at or download app via
Give it a try. Feedback is extremely welcome.
Give it a try. Feedback is extremely welcome.
Introducing Muse, the personal agent that understands your goals and works 24/7 to get things done for you.
It has been long coming, another step towards personal empowerment. You can have a secure, capable and affordable personal agent at your fingertips.
1/ i’m really impressed by the careful work we did to make Muse secure. you can read our technical post here:
thanks for the feedback Kun.
i forced myself to use a few non-mainstream models today. sharing my experience with everyone in case you're wondering 1. muse spark 1.3 i was genuinely surprised by how good this model is. if you silently swapped opus 4.8 with this without telling me, it'll probably take me quite a while to figure it out, and it would likely be from the communication style rather than capability at its contributor pricing, the ROI is quite incredible. i find very little motivation to even try deepseek v4 flash when this exists i don't like the phrase they keep using though - "intelligence that's too cheap to meter". if i do all my work with this model, even at its extremely good pricing, it'll still cost me thousands of dollars a month - that's not "too cheap to meter" reduce the cost by another 100x then let's talk about "too cheap" 2. glm 5.3 flash i thought it'd be better than muse spark 1.3, but actually immediately after i switched to this as my firstmate, it made quite a few mistakes it could be a bit anecdotal but now it lost trust with me and i'm a bit nervous about letting it manage my work. i'm going to let it do some more straightforward implementation rather than acting as my primary model 3. gemini 3.8 flash maybe it's because i never spent much time with gemini before, but this model is SO GOOD at communicating. it's such a breath of fresh air. i can understand every word without using much brain power at all but this is significantly more expensive than the two above. i think i'd have to get a google ai subscription if i want to use this more, but not being able to use it in 3rd party harnesses gives me a pause it's also powerful enough as an opus replacement for most things i tried. my general sense after trying these models is - everyone has an opus now, and they are all cheaper than the real opus fable and astra is probably the only moat the labs have now - the game has fundamentally shifted from what it was 6 months ago
Show more
🥑Muse Spark 1.3 max is a strong model! Close to frontier performance, while sitting on the efficiency frontier. We put a lot of care into post-training it. A good reminder that strong fundamentals, solid execution and attention to detail can take you pretty far. Give it a try — feedback is very welcome!
Show more
updated artificial analysis index—muse spark 1.3 max still performs quite well! the efficient frontier is all Muse, Claude, and GPT
muse spark 1.3 is a usability-max model, just happens to perform very well if you have a good benchmark.
1/ we just publicly released Muse Spark 1.3 max! we see significantly stronger coding and agentic performance on muse spark 1.3 max, so would strongly recommend trying it out even if you've already tried muse spark 1.3 high or muse spark 1.3 xhigh.
Show more
0
142
2K
168
Forward to community
Muse Image is really fun to play with.
Meta's Muse Image debuts at #4# on the Artificial Analysis Image Editing Leaderboard and takes #5# in Text to Image, including a position on the Pareto frontier for quality vs price Muse Image is the first image model from Meta Superintelligence Labs, launched in Meta AI in July and now available to developers on the Meta Model API. Meta positions it as an agentic image model: it invokes search and coding tools, self-refines its own generations, and composes from multiple references. In the Artificial Analysis Image Arena, Muse Image debuts at #4# in Image Editing, behind Microsoft's MAI-Image-2.6-Preview and OpenAI's GPT Image 2 (high) and narrowly behind Google's Nano Banana 2. In Text to Image it takes #5#, behind only GPT Image 2 (high), MAI-Image-2.6-Preview, Reve 2.1, and Google's Nano Banana 2. Muse Image is available in the Meta AI app, on in Instagram Stories, in WhatsApp in limited countries, and to developers via the Meta Model API and partner platforms fal, Runway, and OpenRouter. Congratulations to @AIatMeta, @alexandr_wang, & @finkd on the release! See below for our analysis and example outputs of Muse Image in the Artificial Analysis Image Arena 🧵
Show more
To understand whether we're making genuine progress on reasoning, we entered our AI models in five international STEM Olympiad competitions this year. The results: 🏅 Asian Physics Olympiad (APhO): Perfect score on the theory exam — gold medal 🏅 International Physics Olympiad (IPhO): Perfect score on the theory exam — gold medal 🥇 International Mathematical Olympiad (IMO): Gold medal, top 4% of human participants 🥇 International Chemistry Olympiad (IChO): Gold-medal level performance 🥇 Romanian Masters of Mathematics (RMM): Gold-medal level performance Three of these (APhO, IPhO, IMO) were live competitions and our solutions were submitted under real competition conditions and graded by the official judges using the same marking criteria applied to student contestants. A few things about the approach: • Models were internally trained versions from the Muse Spark family • Zero tool use: no search, no code interpreter, no calculator • Multi-agent orchestration with parallel reasoning We are excited about where this reasoning capability goes next; frontier research level across scientific domains and personal superintelligence. Super grateful to the organizing committees of APhO, IPhO, and IMO for supporting our live participation. We have deep respect for the contestants and organizers behind these competitions. 🙏 And proud of the MSL team that pulled this together!
Show more