가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
가입 January 2009
756 팔로잉 중    43K
Stop everything, benchmark Grok 4.6. And yeah, quite a few steps to get these benchmarks setup. I typically use Fable 5 Low effort as my orchestrator for these. As you can see, lots of updates need to be made to make sure pricing and effort are benchmarked correctly when a new model comes out.
더 보기