# Hugging Face — Deploy Meta Llama 3.1 405B on Google Cloud Vertex AI

- Company: Hugging Face (huggingface.co)
- Announced: 2024-08-19T00:00:00+00:00
- Subject: Platform
- Models affected: meta-llama/Meta-Llama-3.1-405B-Instruct-FP8, meta-llama/Meta-Llama-3.1-405B-Instruct
- Context window: 128K tokens
- Source: https://huggingface.co/blog/llama31-on-vertex-ai
- Record: https://forck.live/items/1837-deploy-meta-llama-3-1-405b-on-google-cloud-vertex-ai

This blog post provides a step-by-step guide on deploying the Meta Llama 3.1 405B model (FP8 quantized variant) on Google Cloud Vertex AI using Hugging Face Deep Learning Containers and Text Generation Inference.

## Evidence

Verbatim from https://huggingface.co/blog/llama31-on-vertex-ai:

> In this blog you will learn how to programmatically deploy meta-llama/Meta-Llama-3.1-405B-Instruct-FP8, the FP8 quantized variant of meta-llama/Meta-Llama-3.1-405B-Instruct, in a Google Cloud A3 node with 8 x H100 NVIDIA GPUs on Vertex AI with Text Generation Inference (TGI) using the Hugging Face purpose-built Deep Learning Containers (DLCs) for Google Cloud.

---

Record: https://forck.live/items/1837-deploy-meta-llama-3-1-405b-on-google-cloud-vertex-ai
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
