Register and share your invite link to earn from video plays and referrals.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
Joined March 2020
63 Following    2.3K Followers
Phew, our longest benchmark run ever is finally done. Muse Spark 1.3 has now been run across every effort level on our v4 eval suite. Will be doing Max as well, but after waiting two weeks for this benchmark to finish, need to take a breather and give some other models some love, then back to Muse. Right now, I think it's safe to say, Muse Spark 1.3 has some serious over-thinking problems at lower effort levels, and Astra is still the strongest for accurate/token efficiency.
Show more
Okay, finally finished my Muse Spark 1.3 benchmark. This ended up being the longest-running benchmark I've ever done at @VulcanBench and it definitely looks like Muse has some serious issues at lower effort levels. I was able to run Astra through my v4 eval suite across every effort level in ~12 hours, Muse Spark had single tasks at low effort levels that hit the 10 hour mark, on just a single task. It took a total of two weeks for me to run the entire benchmark. Overall my assessment is that Muse is a promising model, but they'll have to figure out why it has issues at lower effort levels, looking at the traces, it's just thinking, thinking, and thinking some more, while models like Astra and Fable just got it done. Astra continues to be the most token efficient of these three, and I continue to feel very good about Astra Light, both token efficient and accurate. Using Muse Spark 1.3 at lower effort levels is a complete non-starter imo because you'll be waiting 10 hours for something Astra light can do in 10 minutes. Model cards below, adding to the site soon:
Show more