This is very true. Yes, the classifiers have changed. Hence, it can fallback to Opus more often but that doesn't change the performance of Fable at all. It's just the routing that can be weird and lead to lower "un-supervised" and untrue benchmark results.
Though, Sonnet 5 is bad.
GLM-5.2 on KingBench (3).
Thoughts: The model has superb taste. It is greater at UX than UI. The code is always very clean. It is great at One-shot wonders. I asked it to fine-tune a whole local model and it did it in 30mins!
This is just a great model to use all-round.
1/n