Muse Spark 1.2 is a frontier model for game development. It matches GPT-5.6-Sol, tying for third place.
Earlier this year, Meta didn’t even have a model capable of making this leaderboard. A huge leap for the team—congrats on an amazing model!
Some big updates to GameDevBench!
We’ve tripled (!) the benchmark to 333 tasks and updated the paper accordingly.
We also added GLM-5.2 and Opus 4.8 to the leaderboard.
Opus now leads by a small margin, while GLM-5.2 is the strongest open-source model. This is despite not having any visual capabilities!
Agentic game development remains far from solved, but it’s also exciting to see more agentic game dev work starting to emerge. I’m hoping we’ll see rapid progress here soon.