# Hugging Face — Accelerate Large Model Training using PyTorch Fully Sharded Data Parallel

- Company: Hugging Face (huggingface.co)
- Announced: 2022-05-02
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/pytorch-fsdp
- Record: https://forck.live/items/2207-accelerate-large-model-training-using-pytorch-fully-sharded-data-parallel
- Subject: Platform
- Models affected: GPT-2 Large (762M), GPT-2 XL (1.5B)

To use PyTorch's FullyShardedDataParallel (FSDP) with the Accelerate library to train large models, demonstrating with GPT-2 Large and GPT-2 XL that FSDP enables larger batch sizes and avoids out-of-memory errors compared to Distributed Data Parallel.

## Evidence

Verbatim from https://huggingface.co/blog/pytorch-fsdp:

> In this post we will look at how we can leverage Accelerate Library for training large models which enables users to leverage the latest features of PyTorch FullyShardedDataParallel (FSDP).

---

Record: https://forck.live/items/2207-accelerate-large-model-training-using-pytorch-fully-sharded-data-parallel
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
