# Hugging Face — Llama 2 on Amazon SageMaker a Benchmark

- Company: Hugging Face (huggingface.co)
- Announced: 2023-09-26
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/llama-sagemaker-benchmark
- Record: https://forck.live/items/1991-llama-2-on-amazon-sagemaker-a-benchmark
- Subject: Platform
- Models affected: Llama 2, Llama 2 7B, Llama 2 13B, Llama 2 70B

Hugging Face published a benchmark of Llama 2 model sizes (7B, 13B, 70B) deployed on Amazon SageMaker using the Hugging Face LLM Inference Container, comparing latency and throughput across different instance types, quantization (GPTQ), and concurrency levels to provide recommendations on cost-effective, high-throughput, and low-latency deployments.

## Evidence

Verbatim from https://huggingface.co/blog/llama-sagemaker-benchmark:

> we created a comprehensive benchmark analyzing over 60 different deployment configurations for Llama 2.

---

Record: https://forck.live/items/1991-llama-2-on-amazon-sagemaker-a-benchmark
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
