# Hugging Face — Fine-tune Llama 2 with DPO

- Company: Hugging Face (huggingface.co)
- Announced: 2023-08-08
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/dpo-trl
- Record: https://forck.live/items/2013-fine-tune-llama-2-with-dpo
- Subject: Platform
- Models affected: Llama 2, Llama v2 7B

The Direct Preference Optimization (DPO) method available in the TRL library and demonstrates how to fine-tune the Llama v2 7B model on the stack-exchange preference dataset.

## Evidence

Verbatim from https://huggingface.co/blog/dpo-trl:

> This blog-post introduces the Direct Preference Optimization (DPO) method which is now available in the TRL library and shows how one can fine tune the recent Llama v2 7B-parameter model on the stack-exchange preference dataset

---

Record: https://forck.live/items/2013-fine-tune-llama-2-with-dpo
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
