Model complexity and the KV cache are two primary drivers for memory, especially for inference requiring increasingly large context windows.
Context windows for OpenAI’s leading models have risen 230-260X in three years to 1.05M tokens, while Meta’s Llama 4 Scout is 10X higher at 10M.
$MSFT $AMZN $META $GOOG $MU
Models across the AI ecosystem run on NVIDIA.
New models and AI labs are emerging all the time. Our CEO @JensenHuang shares the idea behind NVIDIA’s platform: build what the ecosystem needs and help everyone succeed. #AllInSummit#
Models bring the intelligence. The Enterprise AI Harness brings the context, action, and control to put it to work.
✨ Highlights from the Enterprise AI Harness Keynote at @Dreamforce with @saastr CEO @jasonlkn and Salesforce product leaders