From the source
Hugging Face and
AMD announce out-of-the-box support for running Transformers models on AMD Instinct GPUs without code changes, including Flash Attention v2, Paged Attention, DeepSpeed, GPTQ, Optimum-Benchmark, and ONNX Runtime integration.
Performance benchmarks show the MI250 GPU delivering over 2.33x more decode throughput and half the prefill latency compared to an A100 card.





