가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Brett Harrison
@BrettHarrison
Founder & CEO @Architect_Fi | Derivatives exchange group for AI commodities and perpetual futures. Offering the American Innovation Exchange and AX.
가입 May 2021
3.3K 팔로잉 중    70.5K 팬
Can inference costs be commoditized? On OpenRouter there’s an 11x spread between the cheapest and most expensive provider of a single model, DeepSeek V4 Flash. Other relevant data points on open-weight inference prices, speeds, availability: • Baidu serves DeepSeek V4 Flash at $0.049 per million tokens and at 124 tokens per second. 26 of the 30 other providers are both more expensive and slower, so it’s not a tradeoff of speed vs cost. • Across the 117 models on OpenRouter with three or more providers, the median price spread between the cheapest and priciest host is 2.2x, the most extreme is 14.4x (DeepSeek v3.2). • Among all endpoints serving the above, 37% of endpoints are strictly dominated by cost, speed, and uptime. The number goes up to 52% if considering only cost and speed.
더 보기