most people forget there are two vectors to optimize for to reduce model cost:
1) reducing the input/output costs, increasing cache hit rates, batching etc
2) packing more intelligence per token (fewer tokens for same task)
inference cost = price per token ร tokens per task
optimizations around the second is underrated and something weโll only see more of.
compression and concision is intelligence.