๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
๊ฐ€์ž… January 2009
756 ํŒ”๋กœ์ž‰ ์ค‘    43K ํŒฌ
Okay, my new 23-task eval suite for @VulcanBench is done. A lot of little details I wanted to get right with this. It was important to me that I had more broad language coverage, and also that it takes into account the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes. I'm new to benchmarking, but learning more every time I create a new set of evals. I'm calling this set v3, getting ready to run the first smoke tests. Then this weekend, phew, safe to say I have a lot of models to test! I will be committing these to Github so anyone can take a look and give me feedback on these as well. Live long and Benchmark ๐Ÿ––
๋” ๋ณด๊ธฐ