# Hugging Face — Direct Preference Optimization Beyond Chatbots

- Company: Hugging Face (huggingface.co)
- Announced: 2026-06-03T12:55:11+00:00
- Category: research-paper
- Subject: Platform
- Models affected: DharmaOCR, Nanonets-OCR2–3B
- Source: https://huggingface.co/blog/Dharma-AI/direct-preference-optimization-beyond-chatbots
- Record: https://forck.live/items/1491-direct-preference-optimization-beyond-chatbots

The post describes using Direct Preference Optimization (DPO) on OCR models to reduce text degeneration, using rejection pairs from the model's own failures. It reports an average reduction of 59.4% in degeneration rate across model families.

## Evidence

Verbatim from https://huggingface.co/blog/Dharma-AI/direct-preference-optimization-beyond-chatbots:

> A second training stage - applied after supervised fine-tuning (SFT), on the same documents, using the same model - reduced text degeneration in every family tested. No exceptions. Average reduction: 59.4%.

---

Record: https://forck.live/items/1491-direct-preference-optimization-beyond-chatbots
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
