๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Yifan Wu
@yifannnwu
ๅดๅฅ•ๅ‡ก; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
๊ฐ€์ž… November 2016
490 ํŒ”๋กœ์ž‰ ์ค‘    1.4K ํŒฌ
RL training for long-horizon tasks is still mysterious, and we took a baby step forward! ๐Ÿง‘โ€๐Ÿณ
Excited to share that SandMLE has been accepted by #COLM2026#! We introduced a multi-agent framework that generates diverse synthetic MLE environments to enable the large-scale on-policy RL training. See you then in SF!
๋” ๋ณด๊ธฐ