@dee_bosa, great point.
This has actually been a thing for the past several months. First, public benchmarks and evals tend to be things that don’t guarantee great results. But more importantly, creating the right evals and driving the model to operate well under those conditions seems like only the normal thing to do.
For example, the EZDubs team was given a very hard time pre-acquisition by
@Cisco for not being too tied to public benchmarks because they wanted to make the right tradeoffs in the model, because the wrong tradeoffs could get you a rank on the benchmark but not allow you to accomplish the results. Specifically, they didn’t want to trade off latency for accuracy. But to score well on the public evals at the time, that would be a necessary trade-off.
So they created their own private evals because it was available in the market was for two generic for them. This will become a more common phenomenon.