I posted this because we need to focus on making real benchmarks of how these models perform in real-world scenarios on real data to evaluate their efficacy.
Just like the factory is really the product, the ability to make proper evals and benchmarks is actually the important thing.
It's too easy to be distracted by the gamed public benchmarks.