# Hugging Face — Why we’re switching to Hugging Face Inference Endpoints, and maybe you should too

- Company: Hugging Face (huggingface.co)
- Announced: 2023-02-15
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/mantis-case-study
- Record: https://forck.live/items/2097-why-we-re-switching-to-hugging-face-inference-endpoints-and-maybe-you-should-too
- Subject: Platform
- Models affected: RoBERTa

The author's team switched to Hugging Face Inference Endpoints for deploying ML models, highlighting easier deployment, lower latency (twice as fast as their previous ECS setup), and a cost increase of 24-50% which they consider acceptable for the time saved.

## Evidence

Verbatim from https://huggingface.co/blog/mantis-case-study:

> Inference Endpoints are more expensive that what we were doing before, there’s an increased cost of between 24% and 50%. At the scale we’re currently operating, this additional cost, a difference of ~$60 a month for a large CPU instance is nothing compared to the time and cognitive load we are saving by not having to worry about APIs, and containers.

---

Record: https://forck.live/items/2097-why-we-re-switching-to-hugging-face-inference-endpoints-and-maybe-you-should-too
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
