When you look at open-weight models processed 56% of tokens on
@vercel's AI Gateway in August, up from 13% in April. That’s 4.3x the share in four months.
When you look even deeper you'll find they only accounted for 14% of estimated spend, while Anthropic still accounted for 64%.
Seems like there’s room for a lot of cheap inference alongside models that can justify a premium.
If more production workloads can run on the cheapest model that clears the quality threshold, application companies could keep more of the value they create.
How much work becomes economical once inference costs stop being the constraint?