๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
๊ฐ€์ž… January 2009
756 ํŒ”๋กœ์ž‰ ์ค‘    43K ํŒฌ
Smoketest passed! Full benchmark running now ๐Ÿ––
Okay, the benchmark I've been waiting all week to run is now running. Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5 Low and Medium effort levels. All on the new v3 eval suite. Can't wait to wake up in the morning to see the results. Live long and benchmark ๐Ÿ––
๋” ๋ณด๊ธฐ