From the source
Hugging Face published a benchmark of Llama 2 model sizes (7B, 13B, 70B) deployed on Amazon SageMaker using the Hugging Face LLM Inference Container, comparing latency and throughput across different instance types, quantization (GPTQ), and concurrency levels to provide recommendations on cost-effective, high-throughput, and low-latency deployments.






