Okay well I finished my Devin SWE-2 results with
@VulcanBench today, so figured I would share them now since I'm already spinning up benchmarks of all the new stuff both OpenAI and Anthropic released today.
It's a busy time to be a benchmarker. And an expensive time to be an independent benchmarker 😅
Overall Devin SWE-2 is a pretty impressive model for the cost, which is $0 right now so pretty hard to beat.
It came in noticeably more accurate than Muse Spark 1.3 and only a few points behind Astra and Fable. So while I might not use it for my hardest of hard tasks, for normal routine tasks I think it's very likely that Devin SWE-2 is more than powerful enough.
Still hard to beat Astra when it comes to speed, it is noticeably faster than Fable, Muse, and Devin. Both Devin and Muse take longer than Fable, they can think a lot on my Frontier v4 suite, which is designed to give them quite a challenge.
I will be running all of these on my new Routine v1 suite which will be pretty interesting as I don't think Muse or Devin are really designed to throw your hardest tasks at, so it will be interesting to see how they do on more routine tasks that Astra and Fable, and now Opus 5.5, are all likely overkill for.
More to come, as always, live long and benchmark 🖖