註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Shobhita Sundaram
@shobsund
PhD student at @MIT_CSAIL working on self-improvement and synthetic data || prev at @GoogleDeepMind @AIatMeta
加入 August 2022
435 正在關注    1.8K 粉絲
Can a model learn to break its own reasoning plateau? In our new paper, we show that LLMs can be taught with meta-RL to generate their own "stepping stones" that kickstart learning on hard math problems (0/128 success rate) where direct RL fails. Paper 📝: Blog post 🌐: (1/n)
顯示更多
0
20
699
114
轉發到社區