Register and share your invite link to earn from video plays and referrals.

Morgan
@morganlinton
cofounder + cto @boldmetrics, eval artist @vulcanbench / prev: @sonos, @carnegiemellon / winter lover, let it snow ❄️
Joined January 2009
870 Following    44.8K Followers
Okay, all my initial benchmark tests with Fable 5.1 in @VulcanBench worked, so I can now run the full sweep. As someone that benchmarks models across all effort levels, and that just raised the timeout from 2 hours to 10 hours, yeah, this is going to take a while. But it's the only way I think we can really do this right and hopefully answer the question: as an engineer, or engineering leader, trying to make sense of when to use the latest frontier model for coding, and at what effort level, what should I do. I hope to help you answer that questions...in 90 - 120 hours.
Show more