註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
加入 January 2009
756 正在關注    43K 粉絲
Comments like this make my day ☀️
It's inspiring me to get deeper into these eval suites. Thanks for sharing the journey! I'm currently messing around with a skill for optimizing orchestration recommendations for model/thinking combinations. Thinking of an ensemble approach where I scan the latest results from DeepSWE—and I'm guessing Vulcan now after glancing at it—for task pass rate, cost, and tokens. Then I have a local eval that measures predicted vs. actual with user survey when PR is done to calibrate over time. Based on X chatter I'm guessing me and thousands of others are trying to do this. It's fun.
顯示更多