HIPPO targets a quiet contaminant in RL for reasoning: pre-RL data overlap.
When the RL dataset overlaps with pretraining or SFT corpora, the model can exploit the shortcut of recalling a memorized answer and then fabricating post-hoc reasoning to match, so the reward goes up while genuine reasoning does not.
The framework injects hints and uses a pairwise objective designed to break that shortcut, forcing the model to reason toward the answer rather than reverse-engineer a justification for one it already memorized.
To Reason or to Fabricate: Reasoning Without Shortcuts via Hint-Anchored Pairwise Optimization
Paper: