注册并分享邀请链接,可获得视频播放与邀请奖励。

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
加入 February 2026
1.8K 正在关注    13.4K 粉丝
A nice theoretical revisit of OPD. With the same supervision budget, supervising only the first 30% of tokens nearly matches standard OPD, while supervising only the last 30% barely works. Framing OPD as a constrained optimization problem naturally leads to an importance-weighted objective, resulting in IW-OPD, which improves sample efficiency and performance with no extra inference cost. Paper: Blog:
显示更多
0
4
305
38
转发到社区