One common pitfall of ML engineering is over-relying on external benchmarks to judge how good a product is without understanding what makes a good product.
A good eval should capture the actual product experience.
If the engineer who trains the model cannot craft one, they do not understand or care about the problem deeply enough.