注册并分享邀请链接,可获得视频播放与邀请奖励。

Ashwinee Panda
@PandaAshwinee
RL Research @togethercompute, Prev: Postdoc of @tomgoldsteincs, PhD @princeton, @Berkeley_EECS alum
加入 February 2020
654 正在关注    3.9K 粉丝
We can now RL large MoEs with 0 train-infer mismatch! And doing so can improve performance (pictured task: teach Qwen3.6-35B-A3B to play Wordle). Everything is open-source and we did a bunch of ablations. 🧵
显示更多
0
14
417
49
转发到社区