๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

(((ู„()(ู„() 'yoav))))๐Ÿ‘พ
@yoavgo
๊ฐ€์ž… May 2009
2.2K ํŒ”๋กœ์ž‰ ์ค‘    89.6K ํŒฌ
why would cache reads be cheaper (assuming its not only business decision)? smaller cache? distilled model? shallower model? fewer tokens?
Cache reads with Fable 5.1 cost 75% less than Fable 5โ€™s. This reduces the cost of the model in practice by around 25% for typical workloads, and up to 45% for highly agentic ones.
๋” ๋ณด๊ธฐ