Muse Spark 1.2 is better than GPT-5.5 xhigh and only slightly worse than Kimi K3 on our ErdosBench.
We've tested the new model from Meta on 226 research-level math problems and it solved 40 / 226 problems and gave many interesting partial solutions.
Muse Spark 1.2 is a strong entrant: good proof hygiene, high B-grade review yield, no rejected strong claims, but fewer decisive A-grade closures.