# Alibaba — GSPO: Towards Scalable Reinforcement Learning for Language Models

- Company: Alibaba (alibaba.com)
- Announced: 2025-07-27T07:00:00+00:00
- Category: research-paper
- Subject: Qwen
- Source: https://qwenlm.github.io/blog/gspo/
- Record: https://forck.live/items/2284-gspo-towards-scalable-reinforcement-learning-for-language-models

Alibaba proposes the Group Sequence Policy Optimization (GSPO) algorithm to address instability and model collapse issues in existing reinforcement learning algorithms for language models, aiming to enable scalable RL training.

## Evidence

Verbatim from https://qwenlm.github.io/blog/gspo/:

> To enable successful RL scaling, we propose the Group Sequence Policy Optimization (GSPO) algorithm.

---

Record: https://forck.live/items/2284-gspo-towards-scalable-reinforcement-learning-for-language-models
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
