Leaderboard spotlight!
With
@Muse on top of the app store, a special shoutout to
@AIatMeta for Muse Spark 1.3. The
@ScaleAILabs team worked closely with them on evaluations, training data, and testing.
Muse does well on tool use, multi-turn conversation, tutoring, professional reasoning, and end-to-end software engineering. Different tasks, but each asks a model to hold context across long horizon workflows, make judgments under uncertainty, and recover from partial failures.
PRBench is where that's clearest. Its tasks are written by domain experts and scored on rubrics that reward professional utility, not trivia-style correctness. Muse has held the top spot in finance and law since 1.1, and 1.3 extends the lead.