RL Infra time.
- nice dispatch strategy that gets rid of long tail stalls
- router replay from previous checkpoints
- this ends up causing shorter completions to impact early stages of training and off-policy. Their solution is to cap at the dataset level (so datasets with general shorter responses dont influence too much) and a discard scheme
- For offpolicyness, they bound the off-policy ratio and loss masking. so pretty standard here
- When a new checkpoint is updated, KVs and routers are persistent instead of being recomputed
At the final stage they do full vocab OPD on >40 teacher models