Our RL training spans four datacenters across three continents, combining our own GPUs across multiple clusters with additional compute from inference providers like
@FireworksAI_HQ.
Only the trainer needs tight collective communication. Rollout inference was distributed, with engines syncing through compressed weight diffs in object storage.