Okay, finally finished my Muse Spark 1.3 benchmark.
This ended up being the longest-running benchmark I've ever done at
@VulcanBench and it definitely looks like Muse has some serious issues at lower effort levels.
I was able to run Astra through my v4 eval suite across every effort level in ~12 hours, Muse Spark had single tasks at low effort levels that hit the 10 hour mark, on just a single task. It took a total of two weeks for me to run the entire benchmark.
Overall my assessment is that Muse is a promising model, but they'll have to figure out why it has issues at lower effort levels, looking at the traces, it's just thinking, thinking, and thinking some more, while models like Astra and Fable just got it done.
Astra continues to be the most token efficient of these three, and I continue to feel very good about Astra Light, both token efficient and accurate. Using Muse Spark 1.3 at lower effort levels is a complete non-starter imo because you'll be waiting 10 hours for something Astra light can do in 10 minutes.
Model cards below, adding to the site soon: