๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
๊ฐ€์ž… March 2020
63 ํŒ”๋กœ์ž‰ ์ค‘    2.3K ํŒฌ
More details on the new code quality score that I'm adding into the new version of VulcanBench. This new frontier is a turning point for benchmarks, excited to be taking the time to make sure we turn in the right direction ๐Ÿ––
๋” ๋ณด๊ธฐ