# Hugging Face — Unlocking Agentic RL Training for GPT-OSS: A Practical Retrospective

- Company: Hugging Face (huggingface.co)
- Announced: 2026-01-27T01:53:15+00:00
- Subject: Platform
- Models affected: GPT-OSS-20B, GPT-OSS-120B, Qwen-2.5-32B
- Source: https://huggingface.co/blog/LinkedIn/gpt-oss-agentic-rl
- Record: https://forck.live/items/1559-unlocking-agentic-rl-training-for-gpt-oss-a-practical-retrospective

The blog post describes a practical retrospective on training the GPT-OSS model with agentic reinforcement learning, detailing challenges such as exploding KL divergence and a fix for a log-probability mismatch in MoE architectures using the verl framework.

## Evidence

Verbatim from https://huggingface.co/blog/LinkedIn/gpt-oss-agentic-rl:

> This blog explores the journey to unlock agentic RL training for GPT-OSS as a potential backbone model for agentic applications.

---

Record: https://forck.live/items/1559-unlocking-agentic-rl-training-for-gpt-oss-a-practical-retrospective
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
