# Alibaba — SAPO: A Stable and Performant Reinforcement Learning Method for Training Large Language Models

- Company: Alibaba (alibaba.com)
- Announced: 2025-12-04T20:00:00+00:00
- Category: research-paper
- Subject: Qwen
- Models affected: Qwen3‑30B‑A3B, Qwen3-30B-A3B-Base
- Source: https://qwen.ai/blog?id=sapo
- Record: https://forck.live/items/4587-sapo-a-stable-and-performant-reinforcement-learning-method-for-training-large

A new reinforcement learning method called SAPO (Soft Adaptive Policy Optimization) is introduced for training large language models, replacing hard clipping with a smooth, temperature-controlled gating function to improve stability and performance.

## Evidence

Verbatim from https://qwen.ai/blog?id=sapo:

> To address these limitations, we propose Soft Adaptive Policy Optimization (SAPO), an RL method designed for stable and performant optimization of LLMs.

---

Record: https://forck.live/items/4587-sapo-a-stable-and-performant-reinforcement-learning-method-for-training-large
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
