From the source
Hugging Face announces integration of 4-bit quantization into transformers using bitsandbytes, enabling running models in 4-bit precision and finetuning with QLoRA, as introduced in the QLoRA paper.
From the source
From the source

Hugging Face announces integration of 4-bit quantization into transformers using bitsandbytes, enabling running models in 4-bit precision and finetuning with QLoRA, as introduced in the QLoRA paper.
From the source
As we strive to make models even more accessible to anyone, we decided to collaborate with bitsandbytes again to allow users to run models in 4-bit precision. This includes a large majority of HF models, in any modality (text, vision, multi-modal, etc.). Users can also train adapters on top of 4bit models leveraging tools from the Hugging Face ecosystem. This is a new method introduced today in the QLoRA paper by Dettmers et al.
huggingface.co