A nice theoretical revisit of OPD.
With the same supervision budget, supervising only the first 30% of tokens nearly matches standard OPD, while supervising only the last 30% barely works.
Framing OPD as a constrained optimization problem naturally leads to an importance-weighted objective, resulting in IW-OPD, which improves sample efficiency and performance with no extra inference cost.
Paper:
Blog: