# Hugging Face — Fine-tuning LLMs to 1.58bit: extreme quantization made easy

- Company: Hugging Face (huggingface.co)
- Announced: 2024-09-18T00:00:00+00:00
- Category: open-weight-release
- Subject: Platform
- Open weights: yes
- Models affected: Llama3 8B, Llama 1 7B, HF1BitLLM/Llama3-8B-1.58-100B-tokens, meta-llama/Meta-Llama-3-8B-Instruct
- Source: https://huggingface.co/blog/1_58_llm_extreme_quantization
- Record: https://forck.live/items/1829-fine-tuning-llms-to-1-58bit-extreme-quantization-made-easy

Hugging Face fine-tuned Llama3 8B models to 1.58-bit quantization using the BitNet architecture, released the models under the HF1BitLLM organization, and introduced a new quantization method called 'bitnet' in Transformers.

## Evidence

Verbatim from https://huggingface.co/blog/1_58_llm_extreme_quantization:

> We have successfully fine-tuned a Llama3 8B model using the BitNet architecture, achieving strong performance on downstream tasks. The 8B models we developed are released under the HF1BitLLM organization.

---

Record: https://forck.live/items/1829-fine-tuning-llms-to-1-58bit-extreme-quantization-made-easy
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
