# Hugging Face — Scaling-up BERT Inference on CPU (Part 1)

- Company: Hugging Face (huggingface.co)
- Announced: 2021-04-20
- Category: research-paper
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/bert-cpu-scaling-part-1
- Record: https://forck.live/items/2259-scaling-up-bert-inference-on-cpu-part-1
- Subject: Platform
- Models affected: BERT

Hugging Face publishes a blog post benchmarking BERT-like model inference on modern CPUs, covering hardware optimizations, core scaling, and batch size scaling using PyTorch, TensorFlow, TorchScript, XLA, and ONNX Runtime on an AWS c5.metal instance with an Intel Xeon Platinum 8275 CPU.

## Evidence

Verbatim from https://huggingface.co/blog/bert-cpu-scaling-part-1:

> This blog post is the first part of a series which will cover most of the hardware and software optimizations to better leverage CPUs for BERT model inference.

---

Record: https://forck.live/items/2259-scaling-up-bert-inference-on-cpu-part-1
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
