On HealthBench Pro, Muse Spark 1.1 achieves similar performance with GPT-5.6 Sol (maybe slightly better) at a fraction of the cost. Affordable health superintelligence is our north star!
We benchmarked Muse Spark 1.1 and GPT-5.6 Sol on HealthBench Professional, OpenAI's benchmark of 525 real clinician tasks 🏥🩺
Muse Spark 1.1 tops our board: better overall score than GPT-5.6 Sol, statistically on par on the length-adjusted score at a fraction of the cost ($1.25/$4.25 vs $5/$30 per M tokens in/out, ~7× cheaper on output).
显示更多