Okay, all my initial benchmark tests with Fable 5.1 in
@VulcanBench worked, so I can now run the full sweep.
As someone that benchmarks models across all effort levels, and that just raised the timeout from 2 hours to 10 hours, yeah, this is going to take a while.
But it's the only way I think we can really do this right and hopefully answer the question: as an engineer, or engineering leader, trying to make sense of when to use the latest frontier model for coding, and at what effort level, what should I do.
I hope to help you answer that questions...in 90 - 120 hours.