Model complexity and the KV cache are two primary drivers for memory, especially for inference requiring increasingly large context windows.
Context windows for OpenAI’s leading models have risen 230-260X in three years to 1.05M tokens, while Meta’s Llama 4 Scout is 10X higher at 10M.
$MSFT $AMZN $META $GOOG $MU
顯示更多