๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Shobhita Sundaram
@shobsund
PhD student at @MIT_CSAIL working on self-improvement and synthetic data || prev at @GoogleDeepMind @AIatMeta
๊ฐ€์ž… August 2022
435 ํŒ”๋กœ์ž‰ ์ค‘    1.8K ํŒฌ
Can a model learn to break its own reasoning plateau? In our new paper, we show that LLMs can be taught with meta-RL to generate their own "stepping stones" that kickstart learning on hard math problems (0/128 success rate) where direct RL fails. Paper ๐Ÿ“: Blog post ๐ŸŒ: (1/n)
๋” ๋ณด๊ธฐ