# Hugging Face — Serverless Inference with Hugging Face and NVIDIA NIM

- Company: Hugging Face (huggingface.co)
- Announced: 2024-07-29T00:00:00+00:00
- Category: infrastructure-release
- Subject: Platform
- Models affected: meta-llama/Meta-Llama-3-8B-Instruct, meta-llama/Meta-Llama-3-70B-Instruct, meta-llama/Meta-Llama-3.1-405B-Instruct-FP8, mistralai/Mixtral-8x22B-Instruct-v0.1, mistralai/Mixtral-8x7B-Instruct-v0.1, mistralai/Mistral-7B-Instruct-v0.3, meta-llama/Meta-Llama-3.1-70B-Instruct, meta-llama/Meta-Llama-3.1-8B-Instruct
- Pricing: $8.25 per hour
- Source: https://huggingface.co/blog/inference-dgx-cloud
- Record: https://forck.live/items/1847-serverless-inference-with-hugging-face-and-nvidia-nim

Hugging Face and NVIDIA launch a serverless inference service called NVIDIA NIM API (serverless) on the Hugging Face Hub, available to Enterprise Hub organizations, providing pay-as-you-go access to open models like Llama and Mistral using NVIDIA DGX Cloud accelerated compute.

## Evidence

Verbatim from https://huggingface.co/blog/inference-dgx-cloud:

> Today, we are thrilled to announce the launch of Hugging Face NVIDIA NIM API (serverless), a new service on the Hugging Face Hub, available to Enterprise Hub organizations.

---

Record: https://forck.live/items/1847-serverless-inference-with-hugging-face-and-nvidia-nim
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
