註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ariel
@ArielKwiat
p/hd | Big RL energy | RS @ big company (not speaking for the company though) | Prev. {Meta FAIR; Gym(nasium)} | Glory to Mankind
加入 November 2011
302 正在關注    6.1K 粉絲
Can we train LLMs with RL using the same next token prediction loss as pre-training? (yes) We conduct a study on (log)prob rewards and show they give a simple way to bridge verifiable and non-verifiable settings with a single reward, broadly applicable for fine-tuning LLMs.
顯示更多
0
5
161
21
轉發到社區