From the source
Runway researchers introduced latent diffusion models (LDMs) that apply diffusion in the latent space of pretrained autoencoders, using cross-attention layers for conditioning inputs such as text or bounding boxes.
LDMs achieve competitive performance on unconditional image generation, inpainting, and super-resolution while reducing computational requirements compared to pixel-based diffusion models.





