why would cache reads be cheaper (assuming its not only business decision)?
smaller cache? distilled model? shallower model? fewer tokens?
Cache reads with Fable 5.1 cost 75% less than Fable 5’s.
This reduces the cost of the model in practice by around 25% for typical workloads, and up to 45% for highly agentic ones.
顯示更多