Big +1 to this. We were a bit skeptical of this when we were first starting out, but once the task is defined, it's strictly better to move to smaller models. Better latency, better cost, AND you can tune it for the specific task.
Great reason for frontier labs to open-source small models since they don't compete with the big models on the same use cases.
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire.
finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.