๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

VulcanBench
@VulcanBench
Open Source LLM benchmarking tool, focused on real world tests, large codebases, full transparency. An Open Source project by @morganlinton.
๊ฐ€์ž… March 2020
28 ํŒ”๋กœ์ž‰ ์ค‘    842 ํŒฌ
I wanted to share a bit more about the benchmark results of our first voice model comparison with VulcanBench. We asked Grok Voice Think Fast 2.0 and GPT Realtime the same 200 questions, typed vs spoken. Here's the results: Grok: 99.0% text โ†’ 95.7% audio (+3.3 pp) GPT Realtime: 97.5% โ†’ 93.5% (+4.0 pp) Real, but small. And here's a little thread for anyone that wants to dive deeper into the details. A voice model benchmarking thread ๐Ÿงต
๋” ๋ณด๊ธฐ