註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
加入 January 2009
756 正在關注    43K 粉絲
Loading up the APIs for my next @VulcanBench run. Put a lot into this one, v3 eval suite with 23 new evals. Broader language coverage. And some really deep dives into making sure this doesn't have any of the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes. Still, I'm new to benchmarking, so there's always a chance that I waste like $150 here. Hoping that's not the case, feels like I've tried to think of every edge case. This is the comparison I've been so excited to do: Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5 All on Low and Medium effort, across real tasks that represent the work real engineering teams will be doing daily. No puzzles, no weird math problems, and looking not just at accuracy, but token use, cost, and time. In aggregate this will be a total of 138 runs, will wake up on Saturday to see the results.
顯示更多