💭 What if a model reading a long context already knows exactly where to look, but keeps dutifully re-scanning everything anyway?
Long-context models have to sweep their entire KV cache at every decoding step, even though attention actually concentrates on a small handful of tokens. Existing sparse attention methods still need to re-search for those relevant tokens at every step, and that search itself becomes the bottleneck.
Researchers at KAIST AI and Google DeepMind flipped the question: instead of scoring tokens externally, why not let the model declare where it's looking, in its own words? That's Declarative Attention. Inside its chain-of-thought, the model announces one of three modes --
for scanning everything, for a named segment, for just recent output -- and the inference engine turns that declaration directly into an attention mask.
Remarkably, this works zero-shot, with no additional training. On Gemma-4-31B and Qwen-3.6-27B, attended tokens drop by 52.0% and 31.1% respectively, while accuracy only slips by 1-3 points. The effect gets more reliable as models and contexts grow larger, saving up to 21 million tokens per response on the longest cases.
Language Models Can Control Their Own Attention
By making attention patterns auditable while cutting inference cost, this looks like a real lever for lowering the cost of long-horizon agentic reasoning.
#LLM# #Attention#