# Hugging Face — Putting RL back in RLHF

- Company: Hugging Face (huggingface.co)
- Announced: 2024-06-12T00:00:00+00:00
- Category: developer-tool-release
- Subject: Platform
- Source: https://huggingface.co/blog/putting_rl_back_in_rlhf_with_rloo
- Record: https://forck.live/items/1873-putting-rl-back-in-rlhf

Hugging Face introduces the RLOO (REINFORCE Leave One-Out) Trainer in TRL, a new online RLHF training algorithm that is more memory-efficient and faster than PPO.

## Evidence

Verbatim from https://huggingface.co/blog/putting_rl_back_in_rlhf_with_rloo:

> We are excited to introduce the RLOO (REINFORCE Leave One-Out) Trainer in TRL.

---

Record: https://forck.live/items/1873-putting-rl-back-in-rlhf
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
