doesn’t it feel like every AI company is starting to do everything??
neoclouds are adding inference. inference companies are reserving or buying gpus. even routers (like openrouter and vercel ai gateway) need to guarantee gpu capacity to make sure requests get served.
an intuitive reason for this is everyone wants to make more money, but I think it's more nuanced. what everyone wants is to control where tokens go and under that is needing to control the gpu supply
- for neoclouds, they need to maximize the usage of gpus (so they add inference to do this)
- inference companies would never want someone else (like AWS or a neocloud) to be the reason why they can't serve all requests (so they reserve or own gpus so they're more independent)
- routers need compute providers with enough available gpus (or there’s nowhere to send requests)
when it gets harder to get compute (like it is right now) it makes sense for even routers to eventually reserve gpus themselves.
in the next few years, these labels are going to get outdated. I predict that many companies will be doing many of the same things