註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ashwinee Panda
@PandaAshwinee
RL Research @togethercompute, Prev: Postdoc of @tomgoldsteincs, PhD @princeton, @Berkeley_EECS alum
加入 February 2020
654 正在關注    3.9K 粉絲
We can now RL large MoEs with 0 train-infer mismatch! And doing so can improve performance (pictured task: teach Qwen3.6-35B-A3B to play Wordle). Everything is open-source and we did a bunch of ablations. 🧵
顯示更多
0
14
417
49
轉發到社區