The hardest bugs were numerical. Long-horizon RL drifts when the engine generating rollouts and the one scoring them stop assigning the same probabilities to the same tokens. Aligning tokenization across both kept the training signal trustworthy.
顯示更多