Parallel inference is inevitable.
@adityagrover_ took the
@Ai4Conferences stage to make the case:
GPUs parallelized matrix multiplication.
Transformers parallelized training.
Diffusion parallelizes inference.
Sequential token generation has a ceiling. Diffusion breaks it.