From the source
Simple optimizations for Stable Diffusion XL (SDXL) to improve inference speed and reduce memory usage, including lower precision (fp16), memory-efficient attention (SDPA), and torch.compile.
From the source
From the source

Simple optimizations for Stable Diffusion XL (SDXL) to improve inference speed and reduce memory usage, including lower precision (fp16), memory-efficient attention (SDPA), and torch.compile.
From the source
To explore how we can optimize SDXL for inference speed and memory use, we ran some tests on an A100 GPU (40 GB).
huggingface.co