Register and share your invite link to earn from video plays and referrals.

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
Joined February 2026
1.8K Following    13.4K Followers
A nice theoretical revisit of OPD. With the same supervision budget, supervising only the first 30% of tokens nearly matches standard OPD, while supervising only the last 30% barely works. Framing OPD as a constrained optimization problem naturally leads to an importance-weighted objective, resulting in IW-OPD, which improves sample efficiency and performance with no extra inference cost. Paper: Blog:
Show more