登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
参加 February 2026
1.8K フォロー中    13.4K ファン
A nice theoretical revisit of OPD. With the same supervision budget, supervising only the first 30% of tokens nearly matches standard OPD, while supervising only the last 30% barely works. Framing OPD as a constrained optimization problem naturally leads to an importance-weighted objective, resulting in IW-OPD, which improves sample efficiency and performance with no extra inference cost. Paper: Blog:
もっと見る