inference is already making compute tradable. if you’re bearish on trading gpus, think about inference! my friend
@lihanc02 described inference providers as “partly compute traders” !!
they reserve compute, run models on it, then sell the output by the token. they make money by buying compute cheaply and getting more tokens from each gpu, but take the risk that some of the compute will not be busy
they take all different kinds of compute and turns them into something that's MUCH easier to compare (the same model’s tokens at a set price and speed)
you can already see this happening in the inference market:
- turn gpu capacity into tokens:
@togethercompute ,
@FireworksAI_HQ,
@DeepInfra,
@baseten
- turn their own chips into tokens:
@GroqLLC,
@cerebras,
@SambaNovaAI
- compare token price and speed:
@ArtificialAnlys
- compare and route requests between providers:
@OpenRouter,
@Vercel,
@RequestyAI
there’s no open market for inference yet, but lots of the pieces are already here