# Hugging Face — Optimum-NVIDIA Unlocking blazingly fast LLM inference in just 1 line of code

- Company: Hugging Face (huggingface.co)
- Announced: 2023-12-05
- Category: developer-tool-release
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/optimum-nvidia
- Record: https://forck.live/items/1970-optimum-nvidia-unlocking-blazingly-fast-llm-inference-in-just-1-line-of-code
- Subject: Platform
- Models affected: meta-llama/Llama-2-7b-chat-hf, meta-llama/Llama-2-13b-chat-hf

Hugging Face announces Optimum-NVIDIA, an inference library that accelerates LLM inference on NVIDIA GPUs with a simple API change. It claims up to 28x faster inference and 1,200 tokens/second, supports FP8, and provides code examples for LLaMA models.

## Evidence

Verbatim from https://huggingface.co/blog/optimum-nvidia:

> By changing just a single line of code, you can unlock up to 28x faster inference and 1,200 tokens/second on the NVIDIA platform.

---

Record: https://forck.live/items/1970-optimum-nvidia-unlocking-blazingly-fast-llm-inference-in-just-1-line-of-code
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
