From the source
Hugging Face integrated the AutoGPTQ library into Transformers, enabling users to quantize and run models in 8, 4, 3, or 2-bit precision using the GPTQ algorithm with negligible accuracy degradation for 4-bit quantization.
From the source
From the source

Hugging Face integrated the AutoGPTQ library into Transformers, enabling users to quantize and run models in 8, 4, 3, or 2-bit precision using the GPTQ algorithm with negligible accuracy degradation for 4-bit quantization.
From the source
we have just integrated the AutoGPTQ library in Transformers, making it possible for users to quantize and run models in 8, 4, 3, or even 2-bit precision using the GPTQ algorithm (Frantar et al. 2023). There is negligible accuracy degradation with 4-bit quantization, with inference speed comparable to the fp16 baseline for small batch sizes.
huggingface.co