My lab's research is supported by both
@thinkymachines and
@river_ai_inc, and I asked one of my students to compare the baseline and RL (GRPO) performance of GLM 5.2 through both APIs, to see how much they disagree. The task is memory based agentic long-horizon, and tldr the two APIs have very similar performance! I'm quite shocked how close the final model outcomes are.