Experimenting with more ways to visualize my benchmark data with
@VulcanBench
I really want to give real insights into what model and effort level you actually need for routine, daily coding tasks.
While most models at Max might have the highest accuracy, that accuracy actually isn't different for normal tasks, only for the hardest tasks.
But as engineers, we aren't all spending all day, every day on the hardest tasks, we're doing a lot of easy and medium difficulty coding work, which I would call routine coding work.
Curious what people think of something like this to better translate my benchmarks into real decision processes, example below: