๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
๊ฐ€์ž… January 2009
756 ํŒ”๋กœ์ž‰ ์ค‘    43K ํŒฌ
Okay, just saw the @VulcanBench results for Grok 4.6 across all effort levels, and this might be the most interesting benchmark I have ever done. This is one pass, think I'll probably need to do three passes, but I'll share the results in the morning. Kinda need to see if someone at @SpaceXAI might be able to hook me up with some credits though. It'll cost me around $150 for me to do another two passes, which I'll totally do, but feels like I need to see if there's a time to ask, it would be now. So...if anyone knows someone that might have some power to help an independent benchmarking nerd like me, please send them my way ๐Ÿ––
๋” ๋ณด๊ธฐ