๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
๊ฐ€์ž… January 2009
756 ํŒ”๋กœ์ž‰ ์ค‘    43K ํŒฌ
Well it finally happened. A lab gave me credits to use for benchmarking with @VulcanBench so I don't have to move out of my house and into a paper box in order to afford these benchmarks ๐Ÿ˜… ๐Ÿ“ฆ Huge thanks to @cursor_ai, who has now made it possible for me to run a 3 pass benchmark with Grok 4.6, and likely whatever else comes out this week, without breaking the bank ๐Ÿ™ Proud to bring independent benchmarking to the world, with 100% real engineering tasks, the stuff actual engineering teams would do with a model. No puzzles, no exotic architecture. And all open source, so every eval, and the benchmarking software itself is all available for you to review, critique, fork, etc. Live long and benchmark ๐Ÿ––
๋” ๋ณด๊ธฐ