We’re constantly running A/B tests in production with different AI models.
Increasingly, we’re finding that across some of our core use cases, users simply can’t tell the difference between frontier models and models that are 50× cheaper.
And this is happening more and more often.