登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Morgan
@morganlinton
cofounder + cto @boldmetrics, eval artist @vulcanbench
参加 January 2009
878 フォロー中    45K ファン
Okay, all my initial benchmark tests with Fable 5.1 in @VulcanBench worked, so I can now run the full sweep. As someone that benchmarks models across all effort levels, and that just raised the timeout from 2 hours to 10 hours, yeah, this is going to take a while. But it's the only way I think we can really do this right and hopefully answer the question: as an engineer, or engineering leader, trying to make sense of when to use the latest frontier model for coding, and at what effort level, what should I do. I hope to help you answer that questions...in 90 - 120 hours.
もっと見る