Register and share your invite link to earn from video plays and referrals.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
Joined March 2020
63 Following    2.3K Followers
Update on my Muse Spark 1.3 benchmark. TL;DR, I think it's going to be a while.
Muse Spark 1.3 is the slowest model I've benchmarked on VulcanBench so far. I don't quite know what is going on, but it took 51.3 hours, on it's lowest effort level, to complete the 23 tasks in VulcanBench-SWE v4. For comparison, it took Astra 1.6 hours on Low Effort to complete the same 23 tasks. Also only 10/23 full passes so pretty disappointing on the accuracy side. Not sure what's going on with Muse, going through the traces to try to understand this better. I can't continue running the benchmark until the Sept 14th as I hit a usage limit on my $50 Muse Code plan. For comparison, I was able to run every effort level, with Astra, on my $100 plan and still have room to spare. If anyone from Meta wants to look at the traces with me you're welcome to, this is a weird one. At this rate, it might take me a month or longer to benchmark this model. For comparison, exact same full effort sweep took ~12 hours with Astra. For some reason I thought Muse Spark would be faster/more token efficient than Astra...but it's not remotely close.
Show more