登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
参加 February 2026
1.7K フォロー中    12.4K ファン
GEAR turns long-context grounding into a reward-shaping problem, by rewarding n-gram overlap with annotated evidence while penalizing overlap with distractors, reducing both copying and reasoning length. The method still depends on automatically generated evidence annotations and overlap proxies, but it gives long-context RL a concrete selectivity objective instead of another generic accuracy bonus. Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning Paper:
もっと見る