Register and share your invite link to earn from video plays and referrals.

Ashwinee Panda
@PandaAshwinee
RL Research @togethercompute, Prev: Postdoc of @tomgoldsteincs, PhD @princeton, @Berkeley_EECS alum
654 Following    3.9K Followers
We can now RL large MoEs with 0 train-infer mismatch! And doing so can improve performance (pictured task: teach Qwen3.6-35B-A3B to play Wordle). Everything is open-source and we did a bunch of ablations. 🧵
Show more