# Hugging Face — Make your llama generation time fly with AWS Inferentia2

- Company: Hugging Face (huggingface.co)
- Announced: 2023-11-07
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/inferentia-llama2
- Record: https://forck.live/items/1974-make-your-llama-generation-time-fly-with-aws-inferentia2
- Subject: Platform
- Models affected: Llama 2, Llama2 7B, Llama2 13B, meta-llama/Llama-2-7b-hf
- Context window: maximum sequence length of 2048

The blog post from Hugging Face explains how to use optimum-neuron to deploy Llama 2 models for text generation on AWS Inferentia2, covering setup, model export, generation, and benchmarks.

## Evidence

Verbatim from https://huggingface.co/blog/inferentia-llama2:

> In a further step of integration with the AWS Neuron SDK, it is now possible to use 🤗 optimum-neuron to deploy LLM models for text generation on AWS Inferentia2.

---

Record: https://forck.live/items/1974-make-your-llama-generation-time-fly-with-aws-inferentia2
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
