๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Zain
@zainhas
I build and teach AI โ€ข AI/ML @togethercompute โ€ข EngSci โ„•ฮจ/PhD @UofT โ€ข Previously: vector DBs, data scientist, lecturer & health tech founder โ€ข ๐Ÿ‡บ๐Ÿ‡ธ๐Ÿ‡จ๐Ÿ‡ฆ๐Ÿ‡ต๐Ÿ‡ฐ
๊ฐ€์ž… August 2012
2.1K ํŒ”๋กœ์ž‰ ์ค‘    8.3K ํŒฌ
probably the most important thing about this release: "Compared with the previous generation, V4.1-Flashโ€™s KV cache needs just: ๐Ÿ”น 1/4 the HBM ๐Ÿ”น 1/8 the SSD storage" 4x smaller than even v4 flash
๋” ๋ณด๊ธฐ
๐Ÿ’พ Smaller KV cache. Bigger savings. Compared with the previous generation, V4.1-Flashโ€™s KV cache needs just: ๐Ÿ”น 1/4 the HBM ๐Ÿ”น 1/8 the SSD storage Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly. 3/6
๋” ๋ณด๊ธฐ