๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
๊ฐ€์ž… March 2020
63 ํŒ”๋กœ์ž‰ ์ค‘    2.3K ํŒฌ
I am going to be dedicating some time to a new eval suite. We need more benchmarks exploring new ways to measure model safety. As a truly independent benchmark, made by me, just one guy, who has never worked at an AI lab in my life, I feel I can bring a unique perspective. Coming soon ๐Ÿ––
๋” ๋ณด๊ธฐ