# Hugging Face — Incredibly Fast BLOOM Inference with DeepSpeed and Accelerate

- Company: Hugging Face (huggingface.co)
- Announced: 2022-09-16
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/bloom-inference-pytorch-scripts
- Record: https://forck.live/items/2153-incredibly-fast-bloom-inference-with-deepspeed-and-accelerate
- Subject: Platform
- Models affected: BLOOM

To achieve fast per-token inference throughput for the 176B-parameter BLOOM model using DeepSpeed and Accelerate, including benchmarks on 8x80GB A100 GPUs.

## Evidence

Verbatim from https://huggingface.co/blog/bloom-inference-pytorch-scripts:

> This article shows how to get an incredibly fast per token throughput when generating with the 176B parameter BLOOM model.

---

Record: https://forck.live/items/2153-incredibly-fast-bloom-inference-with-deepspeed-and-accelerate
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
