가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
가입 January 2009
756 팔로잉 중    43K
Comments like this make my day ☀️
It's inspiring me to get deeper into these eval suites. Thanks for sharing the journey! I'm currently messing around with a skill for optimizing orchestration recommendations for model/thinking combinations. Thinking of an ensemble approach where I scan the latest results from DeepSWE—and I'm guessing Vulcan now after glancing at it—for task pass rate, cost, and tokens. Then I have a local eval that measures predicted vs. actual with user survey when PR is done to calibrate over time. Based on X chatter I'm guessing me and thousands of others are trying to do this. It's fun.
더 보기