登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Seth Karten
@sethkarten
Prime Agent | Continual Harness, LLM Economist | Research @PrimeIntellect | PhD @Princeton | Former CMU Waymo | NSF GRFP
参加 October 2012
690 フォロー中    3.2K ファン
RTing this because people still don't know that training on games can help improve performance on math and reasoning benchmarks, zero-shot
In our latest paper, we discovered a surprising result: training LLMs with self-play reinforcement learning on zero-sum games (like poker) significantly improves performance on math and reasoning benchmarks, zero-shot. Whaaat? How does this work? We analyze the results and find that LLMs learn emergent reasoning patterns like case-by-case analysis and expected value calculation that transfer to improve performance on math questions. This work shows the benefit of RL training for improving reasoning skills when there is no possibility for data leakage. AND how continuously evolving multi-agent competition leads to the development of emergent skills that generalize to novel tasks. Read more below!
もっと見る