가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
가입 March 2020
63 팔로잉 중    2.3K 팬
Noodling on some different ways to visualize the VulcanBench results, would love feedback, more details below:
Experimenting with more ways to visualize my benchmark data with @VulcanBench I really want to give real insights into what model and effort level you actually need for routine, daily coding tasks. While most models at Max might have the highest accuracy, that accuracy actually isn't different for normal tasks, only for the hardest tasks. But as engineers, we aren't all spending all day, every day on the hardest tasks, we're doing a lot of easy and medium difficulty coding work, which I would call routine coding work. Curious what people think of something like this to better translate my benchmarks into real decision processes, example below:
더 보기