註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Graham Neubig
@gneubig
Associate professor @LTIatCMU. Co-founder/chief scientist @OpenHandsDev. I mostly work on modeling language.
加入 September 2010
801 正在關注    47.1K 粉絲
RL helps models learn how to reason with different strategies, but some strategies are more effective than others. But are the strategies learned by RL the ones that are most effective in improving accuracy? Our new work finds that the answer is "not always"!
顯示更多