Register and share your invite link to earn from video plays and referrals.

VulcanBench
@VulcanBench
Open Source LLM benchmarking tool, focused on real world tests, large codebases, full transparency. An Open Source project by @morganlinton.
Joined March 2020
28 Following    842 Followers
Thanks to Fable, VulcanBench v2 is now live on Github. The goal with this release was a big one, so I wanted to wait for Fable to come back to do it. I kept paying $100+ to run my benchmarks where every frontier model scored 98-100%. So I rebuilt it with Fable. VulcanBench v2 is 10 tasks pulled from real merged PRs in flask, aiohttp, sqlglot, click, and chi (Python + Go). Everything merged after model training cutoffs so nothing is memorized, all graded by deterministic hidden tests. Oh and one really nice addition that seemed like the best way to optimize my costs...prompt caching in the harness, a full two-model run now costs $7.36 instead of $100+. Keep calm and benchmark on 🖖
Show more