# Hugging Face — Illustrating Reinforcement Learning from Human Feedback (RLHF)

- Company: Hugging Face (huggingface.co)
- Announced: 2022-12-09
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://huggingface.co/blog/rlhf
- Record: https://forck.live/items/2122-illustrating-reinforcement-learning-from-human-feedback-rlhf
- Subject: Platform

The process of Reinforcement Learning from Human Feedback (RLHF) in three steps: pretraining a language model, training a reward model with human feedback, and fine-tuning the LM with reinforcement learning.

## Evidence

Verbatim from https://huggingface.co/blog/rlhf:

> In this blog post, we’ll break down the training process into three core steps: Pretraining a language model (LM), gathering data and training a reward model, and fine-tuning the LM with reinforcement learning.

---

Record: https://forck.live/items/2122-illustrating-reinforcement-learning-from-human-feedback-rlhf
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
