The latest version of ArXivLean has been released, which sees a big jump in model performance! GPT-5.6-Sol performs best, solving 18/48, ahead of Opus-5 with 15/48 and Aristotle with 12/48 statements formally proved in Lean.
We have added Kimi K3 and Muse Spark 1.1 to the MathArena leaderboard! Kimi K3 now ranks as the strongest open model at #5#, behind GPT-5.6,5.5 and Fable.
The latest versions of ArXivMath and BrokenArXiv have been released! GPT 5.6 Sol takes top spot on ArXivMath, while GPT 5.5 barely manages to stay on top on BrokenArXiv. Fable 5 is the strongest non-GPT model, reaching second place on ArXivMath and third place on BrokenArXiv.
We just evaluated GLM 5.2 on Matharena!
Although GLM 5.2 has shown to be very good at coding, the improvement is not as drastic for math. GLM 5.2 beats GLM 5.1, its predecessor by only 1.9% in expected performance.
Introducing two new research-level mathematical training datasets!
Training data for research mathematics, especially in the post-training regime, is severely lacking. Using our benchmark pipelines on a larger scale, we now created almost 6,000 training data samples.
Introducing a new paper, my first with the SRI Lab, where I have now officially started my PhD!
Lean proving agents are expensive. We came up with a method to efficiently query these agents, maintaining performance while decreasing query cost by 25.8%.
Introducing a new paper, my first with the SRI Lab, where I have now officially started my PhD!
Lean proving agents are expensive. We came up with a method to efficiently query these agents, maintaining performance while decreasing query cost by 25.8%.