PyTorch 2.13 is here, with 3,328 commits from 526 contributors and updates across FlexAttention, CuTeDSL, nn.LinearCrossEntropyLoss, torchcomms, FSDP2, Python 3.15 wheels, ROCm, Arm, and XPU.
The release blog and notes cover FlexAttention on Apple Silicon with up to ~12x speedup over SDPA on sparse patterns, a deterministic backward path on CUDA, the CuTeDSL "Native DSL" backend for Inductor, nn.LinearCrossEntropyLoss to reduce peak GPU memory by up to 4x, torchcomms for large-cluster training, and FSDP2 communication overlap improvements.
On July 22 at 11 a.m. PT, join
@albanDesmaison (
@Meta), Andrey Talman (
@Meta), Piotr Bialecki (
@NVIDIA), and Chris Gottbrath for a live 2.13 Q&A.
🔗 Read the release blog, and register for the live Q&A: