Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog post provides a step-by-step guide on deploying the Meta Llama 3.1 405B model (FP8 quantized variant) on Google Cloud Vertex AI using Hugging Face Deep Learning Containers and Text Generation Inference.
From the source
In this blog you will learn how to programmatically deploy meta-llama/Meta-Llama-3.1-405B-Instruct-FP8, the FP8 quantized variant of meta-llama/Meta-Llama-3.1-405B-Instruct, in a Google Cloud A3 node with 8 x H100 NVIDIA GPUs on Vertex AI with Text Generation Inference (TGI) using the Hugging Face purpose-built Deep Learning Containers (DLCs) for Google Cloud.
huggingface.co