# Hugging Face — Goodbye cold boot - how we made LoRA Inference 300% faster

- Company: Hugging Face (huggingface.co)
- Announced: 2023-12-05
- Category: capability-change
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/lora-adapters-dynamic-loading
- Record: https://forck.live/items/1971-goodbye-cold-boot-how-we-made-lora-inference-300-faster
- Subject: Platform
- Models affected: Stable Diffusion XL Base 1.0, SDXL

Hugging Face reduced LoRA inference warm-up time from 25s to 3s by dynamically swapping adapters while keeping the base model warm, cutting response time from 35s to 13s and using fewer GPUs.

## Evidence

Verbatim from https://huggingface.co/blog/lora-adapters-dynamic-loading:

> We swap the Stable Diffusion LoRA adapters per user request, while keeping the base model warm allowing fast LoRA inference across multiple users.

---

Record: https://forck.live/items/1971-goodbye-cold-boot-how-we-made-lora-inference-300-faster
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
