๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

VulcanBench
@VulcanBench
Open Source LLM benchmarking tool, focused on real world tests, large codebases, full transparency. An Open Source project by @morganlinton.
๊ฐ€์ž… March 2020
28 ํŒ”๋กœ์ž‰ ์ค‘    842 ํŒฌ
Okay, the benchmark I've been waiting all week to run is now running. Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5 Low and Medium effort levels. All on the new v3 eval suite. Can't wait to wake up in the morning to see the results. Live long and benchmark ๐Ÿ––
๋” ๋ณด๊ธฐ