TL;DR: libp2p TCP transport with dual trainer/syncer roles = chef's kiss
For model training, don't bottleneck and single-point-of -failure yourself with a blob store (S3/R2).
Also, don't bash your forehead against a wall using listening sockets on nodes that may be behind firewalls/NAT/port mapped containers/etc.
Just set up a separate backbone (sync only, non-GPU/training nodes) across the world with super fast WAN and use push only from training/GPU nodes to this layer. Easy peasy, works like a charm.
And bonus, libp2p's TCP transports are reliably better across providers/networks/countries compared to default quic/udp.
Skynet?