Second part of our series on how we built one of the fastest image generation systems.
This time we dive into the diffusion pipeline and show how we hit 0.45s inference with kernel optimizations, quantization-aware distillation, and timestep distillation.