Muse Spark 1.2 is better than GPT-5.5 xhigh and only slightly worse than Kimi K3 on our ErdosBench.
We've tested the new model from Meta on 226 research-level math problems and it solved 40 / 226 problems and gave many interesting partial solutions.
Muse Spark 1.2 is a strong entrant: good proof hygiene, high B-grade review yield, no rejected strong claims, but fewer decisive A-grade closures.
Kimi 2.7 ranked 2nd after Fable 5 and before GPT-5 xhigh
We have re-run our ErdosBench smoke test on 14 problems with Kimi 2.7, Qwen 3.7 Max, Grok 4.3 and compared it with the top performers from previous runs.
Kimi 2.7 is amazingly good. More below.