From the source
Hugging Face's Accelerate library leverages PyTorch's meta device to load and run very large models that do not fit in memory, using techniques like empty model creation and device mapping.
From the source
From the source

Hugging Face's Accelerate library leverages PyTorch's meta device to load and run very large models that do not fit in memory, using techniques like empty model creation and device mapping.
From the source
We'll explain how Accelerate leverages PyTorch features to load and run inference with very large models, even if they don't fit in RAM or one GPU.
huggingface.co