New on the PyTorch Foundation blog:
@AMD and
@Meta contributors share how PyTorch Monarch was brought to AMD Instinct GPUs with ROCm to support fault tolerant distributed training at scale.
The post walks through the ROCm port of Monarch’s GPU runtime and distributed communication stack, then shows how Monarch, TorchFT, and TorchTitan enable healthy replicas to continue training while failed nodes recover and rejoin without a full checkpoint restart.
Validation includes Llama 3 8B training on a 128 GPU AMD Instinct MI300 SLURM cluster and a 256 GPU AMD Instinct MI355 Kubernetes cluster.
Read the full technical deep dive: