easy way to convince yourself this is true is to look at how many "notable benchmarks" from 12mo+ ago saturate at <90%
the frontier models of today should get literally 100% on these kinds things...
it's becoming clearer that all models of the past couple years were trained on lots of bullshit broken data but managed to get good enough recently that you could point them at the data with careful orchestration and review loops and fix most of the issues and now it's kinda fine