Register and share your invite link to earn from video plays and referrals.

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
Joined February 2026
1.8K Following    13K Followers
Criteria-based GRPO + β·OPSD, where the RL objective regularizes self-distillation, avoiding the collapse observed under pure self-distillation 👀
Improving exploration for rubric RL using OPSD to generate guidance for underexplored or penalized criteria. As combining criteria to construct rewards is common maybe it could be useful beyond rubric RL.
Show more