註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Natasha Jaques
@natashajaques
Assistant Professor leading the Social RL Lab @uwcse and Staff Research Scientist at @GoogleAI.
加入 June 2009
1.1K 正在關注    34.2K 粉絲
In our latest paper, we discovered a surprising result: training LLMs with self-play reinforcement learning on zero-sum games (like poker) significantly improves performance on math and reasoning benchmarks, zero-shot. Whaaat? How does this work? We analyze the results and find that LLMs learn emergent reasoning patterns like case-by-case analysis and expected value calculation that transfer to improve performance on math questions. This work shows the benefit of RL training for improving reasoning skills when there is no possibility for data leakage. AND how continuously evolving multi-agent competition leads to the development of emergent skills that generalize to novel tasks. Read more below!
顯示更多
0
6
277
66
轉發到社區