# Hugging Face — AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU

- Company: Hugging Face (huggingface.co)
- Announced: 2023-12-05
- Category: capability-change
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/huggingface-and-optimum-amd
- Record: https://forck.live/items/1969-amd-large-language-models-out-of-the-box-acceleration-with-amd-gpu
- Subject: Platform

Hugging Face and AMD announce out-of-the-box support for running Transformers models on AMD Instinct GPUs without code changes, including Flash Attention v2, Paged Attention, DeepSpeed, GPTQ, Optimum-Benchmark, and ONNX Runtime integration. Performance benchmarks show the MI250 GPU delivering over 2.33x more decode throughput and half the prefill latency compared to an A100 card.

## Evidence

Verbatim from https://huggingface.co/blog/huggingface-and-optimum-amd:

> We now support all Transformers models and tasks on AMD Instinct GPUs. And our collaboration is not stopping here, as we explore out-of-the-box support for diffusers models, and other libraries as well as other AMD GPUs.

---

Record: https://forck.live/items/1969-amd-large-language-models-out-of-the-box-acceleration-with-amd-gpu
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
