# Hugging Face — Case Study: Millisecond Latency using Hugging Face Infinity and modern CPUs

- Company: Hugging Face (huggingface.co)
- Announced: 2022-01-13
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/infinity-cpu-performance
- Record: https://forck.live/items/2230-case-study-millisecond-latency-using-hugging-face-infinity-and-modern-cpus
- Subject: Platform
- Models affected: DistilBERT

Performance benchmarks of Hugging Face Infinity on Intel Xeon CPUs, showing up to 34% better latency/throughput compared to previous generation and up to 800% better than vanilla transformers. It also notes that Infinity is no longer offered as a commercial inference solution as of December 2022.

## Evidence

Verbatim from https://huggingface.co/blog/infinity-cpu-performance:

> This ice-lake optimized Infinity Container can achieve up to 34% better latency & throughput compared to existing cascade-lake-based instances, and up to 800% better latency & throughput compared to vanilla transformers running on ice-lake.

---

Record: https://forck.live/items/2230-case-study-millisecond-latency-using-hugging-face-infinity-and-modern-cpus
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
