# Hugging Face — Optimization story: Bloom inference

- Company: Hugging Face (huggingface.co)
- Announced: 2022-10-12
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/bloom-inference-optimization
- Record: https://forck.live/items/2145-optimization-story-bloom-inference
- Subject: Platform
- Models affected: bloom

Hugging Face published an optimization story detailing how they built an efficient inference server for the BLOOM model, achieving a 5x latency reduction and 50x more throughput.

## Evidence

Verbatim from https://huggingface.co/blog/bloom-inference-optimization:

> This article gives you the behind-the-scenes of how we made an efficient inference server that powers bloom.

---

Record: https://forck.live/items/2145-optimization-story-bloom-inference
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
