Register and share your invite link to earn from video plays and referrals.

Shobhita Sundaram
@shobsund
PhD student at @MIT_CSAIL working on self-improvement and synthetic data || prev at @GoogleDeepMind @AIatMeta
435 Following    1.8K Followers
Can a model learn to break its own reasoning plateau? In our new paper, we show that LLMs can be taught with meta-RL to generate their own "stepping stones" that kickstart learning on hard math problems (0/128 success rate) where direct RL fails. Paper 📝: Blog post 🌐: (1/n)
Show more
0
20
699
114
Forward to community