GEAR turns long-context grounding into a reward-shaping problem, by rewarding n-gram overlap with annotated evidence while penalizing overlap with distractors, reducing both copying and reasoning length.
The method still depends on automatically generated evidence annotations and overlap proxies, but it gives long-context RL a concrete selectivity objective instead of another generic accuracy bonus.
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
Paper: