The River API was tested in this blog post and outperformed Tinker on reinforcement learning runs with identical training code. We spent a lot of effort to get details like routing replay right so you get the best possible results with the API.
We can now RL large MoEs with 0 train-infer mismatch! And doing so can improve performance (pictured task: teach Qwen3.6-35B-A3B to play Wordle). Everything is open-source and we did a bunch of ablations. 🧵