가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

alphaXiv
@askalphaxiv
High fidelity research
가입 November 2023
101 팔로잉 중    56.9K 팬
“Learning to Solve Hard Problems in RL for LLMs by Never Giving Up” This paper shows RL has a Matthew Effect, disproportionately improving problems the model can already solve while barely improving the hardest ones. So they fix this by dynamically reallocating sampling compute, quickly filtering easy prompts while repeatedly sampling hard prompts until a correct solution is found. This lets RL progressively learn from harder problems instead of wasting compute on already-solved ones.
더 보기