注册并分享邀请链接,可获得视频播放与邀请奖励。

Graham Neubig
@gneubig
Associate professor @LTIatCMU. Co-founder/chief scientist @OpenHandsDev. I mostly work on modeling language.
加入 September 2010
801 正在关注    47.1K 粉丝
RL helps models learn how to reason with different strategies, but some strategies are more effective than others. But are the strategies learned by RL the ones that are most effective in improving accuracy? Our new work finds that the answer is "not always"!
显示更多