가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
가입 February 2026
1.8K 팔로잉 중    13.4K
A nice theoretical revisit of OPD. With the same supervision budget, supervising only the first 30% of tokens nearly matches standard OPD, while supervising only the last 30% barely works. Framing OPD as a constrained optimization problem naturally leads to an importance-weighted objective, resulting in IW-OPD, which improves sample efficiency and performance with no extra inference cost. Paper: Blog:
더 보기