Can we train LLMs with RL using the same next token prediction loss as pre-training?
(yes)
We conduct a study on (log)prob rewards and show they give a simple way to bridge verifiable and non-verifiable settings with a single reward, broadly applicable for fine-tuning LLMs.
顯示更多