periodic reminder: cache hit rate and model token efficiency (# of turns, # reasoning tokens) matter more for cost than per-token sticker price
working on more updates to ⬆️ cache hit rate on @FireworksAI_HQ
same model. same list price. 5.7x apart in real cost.
glm-5.2 across 6 hosts:
@FireworksAI_HQ 18% of list
@sferenceai 30%
@tensorx_ai 43%
@nebiustf 100%, caches nothing
the price sheet tells you almost nothing. cache hit rate is the price.