这意味着 1M上下文只需要占用内存 0.89 GB
端侧运行模型未来已至,消费产品的内存瓶颈逐步降低,算力将会成为最大短板。
加上专用计算芯片,未来人人都可以在自己的设备上运行大模型
(虽然运行过程还是要128G,但是至少多Agent占用会极大减少
💾 Smaller KV cache. Bigger savings.
Compared with the previous generation, V4.1-Flash’s KV cache needs just:
🔹 1/4 the HBM
🔹 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.
3/6
显示更多