๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Ashwinee Panda
@PandaAshwinee
RL Research @togethercompute, Prev: Postdoc of @tomgoldsteincs, PhD @princeton, @Berkeley_EECS alum
๊ฐ€์ž… February 2020
654 ํŒ”๋กœ์ž‰ ์ค‘    3.9K ํŒฌ
We can now RL large MoEs with 0 train-infer mismatch! And doing so can improve performance (pictured task: teach Qwen3.6-35B-A3B to play Wordle). Everything is open-source and we did a bunch of ablations. ๐Ÿงต
๋” ๋ณด๊ธฐ