Register and share your invite link to earn from video plays and referrals.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
Joined March 2020
63 Following    2.3K Followers
Sharing as we go given how long our Fable 5.1 benchmark is going to take with our new 10 hour timeouts. Results are in for Low Effort on our new VulcanBench-SWE v4 Eval Suite.
Okay I'm going to continue to share the results of my Fable 5.1 benchmark with @VulcanBench-SWE v4, my brand new eval suite designed to challenge this new class of models that dropped this week. As a reminder, and the reason why my benchmarks take so long to run, I run these across every effort level, and with VulcanBench-SWE v4 I have increased the timeout to 10 hours, up from 2 hours, this means these are a lot more expensive to run, but ensures I'm getting really clean results, not failing because of timeouts. Additionally, I have added more guardrails to prevent cheating and to give partial credit, so a model can still get some credit even if it doesn't ace a task. All of the tasks in my eval suite are real coding tasks, PRs from open source repos, and represent the kind of coding tasks normal engineering teams would be giving to these models. Fable 5.1 Low scored 82.6 with an average time per task of 14 minutes and average cost per task of $4.66. Once all effort levels are complete I will be adding a full model card to the site and detailed PDF on all the runs and data generated from this sweep. Live long and benchmark 🖖
Show more