# Hugging Face — Making LLMs lighter with AutoGPTQ and transformers

- Company: Hugging Face (huggingface.co)
- Announced: 2023-08-23
- Category: developer-tool-release
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/gptq-integration
- Record: https://forck.live/items/2007-making-llms-lighter-with-autogptq-and-transformers
- Subject: Platform

Hugging Face integrated the AutoGPTQ library into Transformers, enabling users to quantize and run models in 8, 4, 3, or 2-bit precision using the GPTQ algorithm with negligible accuracy degradation for 4-bit quantization. The integration supports Nvidia and AMD GPUs.

## Evidence

Verbatim from https://huggingface.co/blog/gptq-integration:

> we have just integrated the AutoGPTQ library in Transformers, making it possible for users to quantize and run models in 8, 4, 3, or even 2-bit precision using the GPTQ algorithm (Frantar et al. 2023). There is negligible accuracy degradation with 4-bit quantization, with inference speed comparable to the fp16 baseline for small batch sizes.

---

Record: https://forck.live/items/2007-making-llms-lighter-with-autogptq-and-transformers
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
