# Hugging Face — Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA

- Company: Hugging Face (huggingface.co)
- Announced: 2023-05-24
- Category: capability-change
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/4bit-transformers-bitsandbytes
- Record: https://forck.live/items/2054-making-llms-even-more-accessible-with-bitsandbytes-4-bit-quantization-and-qlora
- Subject: Platform
- Models affected: LLaMA, T5, Guanaco, GPT-neo-X

Hugging Face announces integration of 4-bit quantization into transformers using bitsandbytes, enabling running models in 4-bit precision and finetuning with QLoRA, as introduced in the QLoRA paper.

## Evidence

Verbatim from https://huggingface.co/blog/4bit-transformers-bitsandbytes:

> As we strive to make models even more accessible to anyone, we decided to collaborate with bitsandbytes again to allow users to run models in 4-bit precision. This includes a large majority of HF models, in any modality (text, vision, multi-modal, etc.). Users can also train adapters on top of 4bit models leveraging tools from the Hugging Face ecosystem. This is a new method introduced today in the QLoRA paper by Dettmers et al.

---

Record: https://forck.live/items/2054-making-llms-even-more-accessible-with-bitsandbytes-4-bit-quantization-and-qlora
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
