From the source
To use PyTorch's FullyShardedDataParallel (FSDP) with the Accelerate library to train large models, demonstrating with GPT-2 Large and GPT-2 XL that FSDP enables larger batch sizes and avoids out-of-memory errors compared to Distributed Data Parallel.





