(1/12) New Paper 📄
Linear-time attentions are reaching frontier scale, their fixed-state memory, however, still saturates as context grows, because all of its capacity is exposed from token 1. Early tokens face no pressure to compress, so they take far more of that capacity than they need. Later tokens inherit a state that is already full.
We introduce incremental memory activation: a fixed-size memory should not expose all of its capacity at once. It should schedule how much is live as the context grows, imposing a bottleneck early that forces the model to summarize rather than memorize, and unlocking fresh capacity later so new information has somewhere to land that isn't already written.
Proteus is the simplest instantiation of that idea. We partition the memory into E blocks and unlock them one at a time as the context advances, gating both reads and writes so locked blocks are neither retrieved from nor updated. The state itself never grows; only the active fraction of it changes. It comes at no extra cost and improves average performance on language modeling and commonsense reasoning (+0.37–1.05 avg. accuracy points). More interestingly, it yields significant gains in long-context handling: up to +8.4 NIAH points at 2× the training context length.
Read more background/method/results below 🧵
顯示更多