I build and teach AI โข AI/ML @togethercompute โข EngSci โฮจ/PhD @UofT โข Previously: vector DBs, data scientist, lecturer & health tech founder โข ๐บ๐ธ๐จ๐ฆ๐ต๐ฐ
probably the most important thing about this release:
"Compared with the previous generation, V4.1-Flashโs KV cache needs just:
๐น 1/4 the HBM
๐น 1/8 the SSD storage"
4x smaller than even v4 flash
๐พ Smaller KV cache. Bigger savings.
Compared with the previous generation, V4.1-Flashโs KV cache needs just:
๐น 1/4 the HBM
๐น 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.
3/6