๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
๊ฐ€์ž… March 2020
63 ํŒ”๋กœ์ž‰ ์ค‘    2.3K ํŒฌ
A lot have people have been asking me why I donโ€™t get early access to all the new models like the cool kids. Easy answer. All the influencers who get early access to models, sharing glowing posts about how much they loved using the models on launch day, often with some example where the models really shines. At VulcanBench, I just share the data, good or bad, and let that be my guide. As long as I continue on this path, I think getting early access isnโ€™t going to happen for me. But thatโ€™s okay, Iโ€™d rather just be able to share the good, the bad, and the ugly, on new models, even if my results come out a week or two after launch. Live long and benchmark ๐Ÿ––
๋” ๋ณด๊ธฐ