# Hugging Face — Finetune Stable Diffusion Models with DDPO via TRL

- Company: Hugging Face (huggingface.co)
- Announced: 2023-09-29
- Category: developer-tool-release
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/trl-ddpo
- Record: https://forck.live/items/1989-finetune-stable-diffusion-models-with-ddpo-via-trl
- Subject: Platform

Hugging Face releases a new DDPOTrainer in the trl library, enabling users to fine-tune Stable Diffusion models with DDPO (Denoising Diffusion Policy Optimization) reinforcement learning.

## Evidence

Verbatim from https://huggingface.co/blog/trl-ddpo:

> we discuss how DDPO came to be, a brief description of how it works, and how DDPO can be incorporated into an RLHF workflow to achieve model outputs more aligned with the human aesthetics. We then quickly switch gears to talk about how you can apply DDPO to your models with the newly integrated DDPOTrainer from the trl library

---

Record: https://forck.live/items/1989-finetune-stable-diffusion-models-with-ddpo-via-trl
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
