Register and share your invite link to earn from video plays and referrals.

Przemek Chojecki | PC
@prz_chojecki
Reasoning Data + Evals for LLMs @ PhD in mathematics.
Joined April 2016
1.1K Following    14.7K Followers
Muse Spark 1.2 is better than GPT-5.5 xhigh and only slightly worse than Kimi K3 on our ErdosBench. We've tested the new model from Meta on 226 research-level math problems and it solved 40 / 226 problems and gave many interesting partial solutions. Muse Spark 1.2 is a strong entrant: good proof hygiene, high B-grade review yield, no rejected strong claims, but fewer decisive A-grade closures.
Show more