Why and how do diffusion models memorize vs generalize? Can we have scaling laws for memorization? This is increasingly relevant scientifically and pragmatically (e.g. Sora 2).
๐จ Our new preprint "On the Edge of Memorization in Diffusion Models" addresses this timely question!