Improving exploration for rubric RL using OPSD to generate guidance for underexplored or penalized criteria. As combining criteria to construct rewards is common maybe it could be useful beyond rubric RL.
OPD for multi-turn scenario ( Teacher determines whether intervention is needed and how confident it is in the supervision, and these are used as weighting factors.