注册并分享邀请链接,可获得视频播放与邀请奖励。

Ariel
@ArielKwiat
p/hd | Big RL energy | RS @ big company (not speaking for the company though) | Prev. {Meta FAIR; Gym(nasium)} | Glory to Mankind
加入 November 2011
302 正在关注    6.1K 粉丝
Can we train LLMs with RL using the same next token prediction loss as pre-training? (yes) We conduct a study on (log)prob rewards and show they give a simple way to bridge verifiable and non-verifiable settings with a single reward, broadly applicable for fine-tuning LLMs.
显示更多
0
5
161
21
转发到社区