To understand whether we're making genuine progress on reasoning, we entered our AI models in five international STEM Olympiad competitions this year.
The results:
🏅 Asian Physics Olympiad (APhO): Perfect score on the theory exam — gold medal
🏅 International Physics Olympiad (IPhO): Perfect score on the theory exam — gold medal
🥇 International Mathematical Olympiad (IMO): Gold medal, top 4% of human participants
🥇 International Chemistry Olympiad (IChO): Gold-medal level performance
🥇 Romanian Masters of Mathematics (RMM): Gold-medal level performance
Three of these (APhO, IPhO, IMO) were live competitions and our solutions were submitted under real competition conditions and graded by the official judges using the same marking criteria applied to student contestants.
A few things about the approach:
• Models were internally trained versions from the Muse Spark family
• Zero tool use: no search, no code interpreter, no calculator
• Multi-agent orchestration with parallel reasoning
We are excited about where this reasoning capability goes next; frontier research level across scientific domains and personal superintelligence.
Super grateful to the organizing committees of APhO, IPhO, and IMO for supporting our live participation. We have deep respect for the contestants and organizers behind these competitions. 🙏
And proud of the MSL team that pulled this together!