註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在關注    56.9K 粉絲
“Learning to Solve Hard Problems in RL for LLMs by Never Giving Up” This paper shows RL has a Matthew Effect, disproportionately improving problems the model can already solve while barely improving the hardest ones. So they fix this by dynamically reallocating sampling compute, quickly filtering easy prompts while repeatedly sampling hard prompts until a correct solution is found. This lets RL progressively learn from harder problems instead of wasting compute on already-solved ones.
顯示更多
0
6
186
15
轉發到社區