Okay I'm going to continue to share the results of my Fable 5.1 benchmark with
@VulcanBench-SWE v4, my brand new eval suite designed to challenge this new class of models that dropped this week.
As a reminder, and the reason why my benchmarks take so long to run, I run these across every effort level, and with VulcanBench-SWE v4 I have increased the timeout to 10 hours, up from 2 hours, this means these are a lot more expensive to run, but ensures I'm getting really clean results, not failing because of timeouts.
Additionally, I have added more guardrails to prevent cheating and to give partial credit, so a model can still get some credit even if it doesn't ace a task.
All of the tasks in my eval suite are real coding tasks, PRs from open source repos, and represent the kind of coding tasks normal engineering teams would be giving to these models.
Fable 5.1 Low scored 82.6 with an average time per task of 14 minutes and average cost per task of $4.66.
Once all effort levels are complete I will be adding a full model card to the site and detailed PDF on all the runs and data generated from this sweep.
Live long and benchmark 🖖