注册并分享邀请链接,可获得视频播放与邀请奖励。

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
加入 March 2020
63 正在关注    2.3K 粉丝
Noodling on some different ways to visualize the VulcanBench results, would love feedback, more details below:
Experimenting with more ways to visualize my benchmark data with @VulcanBench I really want to give real insights into what model and effort level you actually need for routine, daily coding tasks. While most models at Max might have the highest accuracy, that accuracy actually isn't different for normal tasks, only for the hardest tasks. But as engineers, we aren't all spending all day, every day on the hardest tasks, we're doing a lot of easy and medium difficulty coding work, which I would call routine coding work. Curious what people think of something like this to better translate my benchmarks into real decision processes, example below:
显示更多