๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
๊ฐ€์ž… March 2020
63 ํŒ”๋กœ์ž‰ ์ค‘    2.3K ํŒฌ
Cool idea, honored to be included ๐Ÿ––
picking an llm = three leaderboards, three winners. arena vs academic vs whatever dropped this week. weekend project: UnifyBench ( look at which models look strongest, then drill into the underlying benches for the real detail. 451 models, 109 sources;
๋” ๋ณด๊ธฐ