注册并分享邀请链接,可获得视频播放与邀请奖励。

Shobhita Sundaram
@shobsund
PhD student at @MIT_CSAIL working on self-improvement and synthetic data || prev at @GoogleDeepMind @AIatMeta
加入 August 2022
435 正在关注    1.8K 粉丝
Can a model learn to break its own reasoning plateau? In our new paper, we show that LLMs can be taught with meta-RL to generate their own "stepping stones" that kickstart learning on hard math problems (0/128 success rate) where direct RL fails. Paper 📝: Blog post 🌐: (1/n)
显示更多
0
20
699
114
转发到社区