가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Graham Neubig
@gneubig
Associate professor @LTIatCMU. Co-founder/chief scientist @OpenHandsDev. I mostly work on modeling language.
가입 September 2010
801 팔로잉 중    47.1K 팬
RL helps models learn how to reason with different strategies, but some strategies are more effective than others. But are the strategies learned by RL the ones that are most effective in improving accuracy? Our new work finds that the answer is "not always"!
더 보기