# Hugging Face — Deploy GPT-J 6B for inference using Hugging Face Transformers and Amazon SageMaker

- Company: Hugging Face (huggingface.co)
- Announced: 2022-01-11
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/gptj-sagemaker
- Record: https://forck.live/items/2232-deploy-gpt-j-6b-for-inference-using-hugging-face-transformers-and-amazon
- Subject: Platform
- Models affected: GPT-J 6B, GPT-J

Deploying the GPT-J 6B model for inference using Hugging Face Transformers and Amazon SageMaker, focusing on reducing model loading time by using torch.save and torch.load to achieve production-compatible performance.

## Evidence

Verbatim from https://huggingface.co/blog/gptj-sagemaker:

> In this blog post, you will learn how to easily deploy GPT-J using Amazon SageMaker and the Hugging Face Inference Toolkit with a few lines of code for scalable, reliable, and secure real-time inference using a regular size GPU instance with NVIDIA T4 (~500$/m).

---

Record: https://forck.live/items/2232-deploy-gpt-j-6b-for-inference-using-hugging-face-transformers-and-amazon
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
