가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
가입 February 2026
1.7K 팔로잉 중    12.4K
HIPPO targets a quiet contaminant in RL for reasoning: pre-RL data overlap. When the RL dataset overlaps with pretraining or SFT corpora, the model can exploit the shortcut of recalling a memorized answer and then fabricating post-hoc reasoning to match, so the reward goes up while genuine reasoning does not. The framework injects hints and uses a pairwise objective designed to break that shortcut, forcing the model to reason toward the answer rather than reverse-engineer a justification for one it already memorized. To Reason or to Fabricate: Reasoning Without Shortcuts via Hint-Anchored Pairwise Optimization Paper:
더 보기