Can inference costs be commoditized? On OpenRouter there’s an 11x spread between the cheapest and most expensive provider of a single model, DeepSeek V4 Flash. Other relevant data points on open-weight inference prices, speeds, availability:
• Baidu serves DeepSeek V4 Flash at $0.049 per million tokens and at 124 tokens per second. 26 of the 30 other providers are both more expensive and slower, so it’s not a tradeoff of speed vs cost.
• Across the 117 models on OpenRouter with three or more providers, the median price spread between the cheapest and priciest host is 2.2x, the most extreme is 14.4x (DeepSeek v3.2).
• Among all endpoints serving the above, 37% of endpoints are strictly dominated by cost, speed, and uptime. The number goes up to 52% if considering only cost and speed.