가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
가입 February 2026
1.7K 팔로잉 중    12.4K
GEAR turns long-context grounding into a reward-shaping problem, by rewarding n-gram overlap with annotated evidence while penalizing overlap with distractors, reducing both copying and reasoning length. The method still depends on automatically generated evidence annotations and overlap proxies, but it gives long-context RL a concrete selectivity objective instead of another generic accuracy bonus. Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning Paper:
더 보기