In a compute capacity crunch, the ability to create specialized models that are cheaper AND better in quality becomes super valuable.
We've always said that there are 2 ways to cut your inference costs:
1. Squeeze a ton of juice out of inference optimization (this is getting harder and harder)
2. Train your models to be more token efficient and make smaller models better at your task so you are cheaper iso performance.
JUST IN: Startups are increasingly building custom AI models using open-weight alternatives to cut costs & reduce dependence on OpenAI & Anthropic. — Bloomberg