Register and share your invite link to earn from video plays and referrals.

alignedai
@alignedai
applied ai @meta ⁕ cs @ucla ⁕ for the love of the game
4.3K Following    585 Followers
Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on consumer hardware like a Mac or PCs with performant GPUs. In keeping with our long tradition of sharing fundamental AI research, we’re releasing model weights under a permissive Apache 2.0 license. 🧵👇
Show more
0
381
8.3K
1.1K
Forward to community
Today we released Muse Code, a terminal coding agent, powered by our new Muse Spark 1.2 model. Learn more 👇
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
Show more
0
285
730
71
Forward to community
Muse Spark 1.2 by @AIatMeta is in the Agent Arena! Bring your toughest prompts to power the leaderboards. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. In addition to Agent Arena, Muse Spark 1.2 is available in Text, Vision and Code Arena: Frontend.
Show more
@Percivalkatz @alignedai Muse Spark 1.2 is available today in Muse Code and in Meta Model API. App coming soon
Tie for 3rd among US labs for now 👀🫡👀
Meta has released Muse Spark 1.2. It's their third release in four months and scores 54 on the Artificial Analysis Intelligence Index, significantly improving agentic knowledge work capabilities over prior releases and putting Meta next to SpaceXAI in a tie for third place amongst US labs Muse Spark 1.2 (xhigh) lands at 54, up 3 points from Muse Spark 1.1 (51) and 11 points from Muse Spark 1.0 (43, April). It enters effectively tied with GPT-5.5 (xhigh, 55) and Grok 4.5 (high, 54), narrowly behind current frontier models Claude Opus 5 (max, 61), Claude Fable 5 (max w/ fallback, 60), GPT-5.6 Sol (max, 59), and Kimi K3 (max, 57) Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Muse Spark 1.2 gets closer to the frontier on agentic knowledge work. At Muse Spark 1.1's launch, we noted agentic knowledge work as its clearest gap; Muse Spark 1.2's gains help to close this. Its GDPval-AA v2 Elo rose 260 points to 1631, #5# among all models we have benchmarked and ahead of Claude Opus 4.8 (max, 1588). Terminal-Bench 2.1 gained 2 points (78% to 80%), and Tau3-Bench Banking rose 2 points (25% to 27%) ➤ Among the most cost-efficient models at its intelligence level. Muse Spark 1.2 costs $0.40 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing, with only Grok 4.5 (high, $0.37) and GPT-5.6 Sol (medium, $0.39) cheaper in its intelligence cluster - GPT-5.6 Terra (max, $0.51), Kimi K3 (max, $0.86), and GPT-5.5 (xhigh, $1.18) all cost more per task. The cost increase over Muse Spark 1.1 ($0.29 per task) is driven by increased token usage per Intelligence Index task ➤ AA-Omniscience abstention rate increases. The score rose from 18 to 22 as the hallucination rate fell 10 points (38% to 28%) and the attempt rate dropped from 82% to 67%. This heavy abstention (not answering questions when unsure) now drives both the low hallucination rate and a lower accuracy (41% to 38%) ➤ Scientific Reasoning results remain largely unchanged. CritPt notably gained 3 points (15% to 18%), while SciCode fell 2 points (58% to 56%), and Humanity's Last Exam fell 1 point (45% to 44%) Other model details: ➤ Context window: 1M tokens, unchanged from Muse Spark 1.1 ➤ Pricing: unchanged from Muse Spark 1.1: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Availability: Meta's first-party API at launch
Show more
Meta has released Muse Spark 1.2. It's their third release in four months and scores 54 on the Artificial Analysis Intelligence Index, significantly improving agentic knowledge work capabilities over prior releases and putting Meta next to SpaceXAI in a tie for third place amongst US labs Muse Spark 1.2 (xhigh) lands at 54, up 3 points from Muse Spark 1.1 (51) and 11 points from Muse Spark 1.0 (43, April). It enters effectively tied with GPT-5.5 (xhigh, 55) and Grok 4.5 (high, 54), narrowly behind current frontier models Claude Opus 5 (max, 61), Claude Fable 5 (max w/ fallback, 60), GPT-5.6 Sol (max, 59), and Kimi K3 (max, 57) Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Muse Spark 1.2 gets closer to the frontier on agentic knowledge work. At Muse Spark 1.1's launch, we noted agentic knowledge work as its clearest gap; Muse Spark 1.2's gains help to close this. Its GDPval-AA v2 Elo rose 260 points to 1631, #5# among all models we have benchmarked and ahead of Claude Opus 4.8 (max, 1588). Terminal-Bench 2.1 gained 2 points (78% to 80%), and Tau3-Bench Banking rose 2 points (25% to 27%) ➤ Among the most cost-efficient models at its intelligence level. Muse Spark 1.2 costs $0.40 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing, with only Grok 4.5 (high, $0.37) and GPT-5.6 Sol (medium, $0.39) cheaper in its intelligence cluster - GPT-5.6 Terra (max, $0.51), Kimi K3 (max, $0.86), and GPT-5.5 (xhigh, $1.18) all cost more per task. The cost increase over Muse Spark 1.1 ($0.29 per task) is driven by increased token usage per Intelligence Index task ➤ AA-Omniscience abstention rate increases. The score rose from 18 to 22 as the hallucination rate fell 10 points (38% to 28%) and the attempt rate dropped from 82% to 67%. This heavy abstention (not answering questions when unsure) now drives both the low hallucination rate and a lower accuracy (41% to 38%) ➤ Scientific Reasoning results remain largely unchanged. CritPt notably gained 3 points (15% to 18%), while SciCode fell 2 points (58% to 56%), and Humanity's Last Exam fell 1 point (45% to 44%) Other model details: ➤ Context window: 1M tokens, unchanged from Muse Spark 1.1 ➤ Pricing: unchanged from Muse Spark 1.1: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Availability: Meta's first-party API at launch
Show more
0
63
1.2K
117
Forward to community
@AIatMeta needs to share this: muse-spark-1.2 pricing vs muse-spark-1.2-contributor it is much much much cheaper almost free if you are okay with training: great for open source contributors Input: 12.5× Cached input: 75× Output: 21.25× Put another way: Input: 92% cheaper Cached input: 98.67% cheaper Output: 95.29% cheaper
Show more
MS1.2 it’s a good model 😂😂😂
@AIatMeta needs to share this: muse-spark-1.2 pricing vs muse-spark-1.2-contributor it is much much much cheaper almost free if you are okay with training: great for open source contributors Input: 12.5× Cached input: 75× Output: 21.25× Put another way: Input: 92% cheaper Cached input: 98.67% cheaper Output: 95.29% cheaper
Show more
try this out! curl -fsSL | bash
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
Show more
try this out!!!! curl -fsSL | bash
muse code in beta is here: our first coding agent powered by our latest model, muse spark 1.2. one command to install and start building. get it through Meta Model API.
Muse Spark 1.2 from @AIatMeta is live on OpenRouter alongside expanded global access to both Muse Spark models. At $1.25/M in and $4.25/M out, the model builds its position as one of the most price-efficient, high-intelligence models on OpenRouter.
Show more
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
Show more
0
1.4K
14.8K
1.3K
Forward to community
@NASA’s new telescope, Roman, is coming online this month, it will see 200× more sky than @NASAHubble. A survey that would take Hubble a century will take Roman one month. It will see 20 billion stars, discover tens of thousands of planets, map 2 billion galaxies. 🌌🪐🌌
Show more
@NASA’s new telescope, Roman, is coming online this month, it will see 200× more sky than @NASAHubble. A survey that would take Hubble a century will take Roman one month. It will see 20 billion stars, discover tens of thousands of planets, map 2 billion galaxies. 🌌🪐🌌
Show more
I suspect there are multiple instances of benchmarks missing the mark due to some esoteric config issue.
We implemented the harness with the Responses API and turned on: → Retained reasoning → Context compaction On the public set, GPT-5.6 Sol’s score rose 188% while using 6x fewer output tokens.
Show more
My favorite way to start my day is to take my morning coffee and go for a walk around the block. The vibe is immediately set for the day.