注册并分享邀请链接,可获得视频播放与邀请奖励。

Morgan
@morganlinton
cto @boldmetrics, mad scientist @vulcanbench, giving back at
加入 January 2009
729 正在关注    42.8K 粉丝
Loading up the APIs for my next @VulcanBench run. Put a lot into this one, v3 eval suite with 23 new evals. Broader language coverage. And some really deep dives into making sure this doesn't have any of the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes. Still, I'm new to benchmarking, so there's always a chance that I waste like $150 here. Hoping that's not the case, feels like I've tried to think of every edge case. This is the comparison I've been so excited to do: Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5 All on Low and Medium effort, across real tasks that represent the work real engineering teams will be doing daily. No puzzles, no weird math problems, and looking not just at accuracy, but token use, cost, and time. In aggregate this will be a total of 138 runs, will wake up on Saturday to see the results.
显示更多