BofA’s Vivek Arya made an excellent point, which I’d like to share.
Is open-source AI bearish for memory?
Closed models amortize global demand across shared HBM pools concentrated in a handful of data centers, whereas open models create a new memory footprint with every deployment. If 10,000 companies self-host the same open model, the model weights must be replicated across 10,000 separate HBM pools, with each deployment also requiring its own KV cache. As 128K–1M token contexts become commonplace in 2026, the KV cache alone can exceed 40GB per active session.
Low-cost Chinese APIs drive greater inference demand, broader enterprise self-hosting, and more memory sockets worldwide. In short, closed models concentrate memory demand, while open models multiply it.