Register and share your invite link to earn from video plays and referrals.

Ariel
@ArielKwiat
p/hd | Big RL energy | RS @ big company (not speaking for the company though) | Prev. {Meta FAIR; Gym(nasium)} | Glory to Mankind
Joined November 2011
302 Following    6.1K Followers
Can we train LLMs with RL using the same next token prediction loss as pre-training? (yes) We conduct a study on (log)prob rewards and show they give a simple way to bridge verifiable and non-verifiable settings with a single reward, broadly applicable for fine-tuning LLMs.
Show more