# Alibaba — QwQ-32B: Embracing the Power of Reinforcement Learning

- Company: Alibaba (alibaba.com)
- Announced: 2025-03-05T16:00:04+00:00
- Category: research-paper
- Subject: Qwen
- Source: https://qwenlm.github.io/blog/qwq-32b/
- Record: https://forck.live/items/2294-qwq-32b-embracing-the-power-of-reinforcement-learning

Alibaba's Qwen team discusses the potential of reinforcement learning to improve model performance, citing examples like DeepSeek R1.

## Evidence

Verbatim from https://qwenlm.github.io/blog/qwq-32b/:

> Scaling Reinforcement Learning (RL) has the potential to enhance model performance beyond conventional pretraining and post-training methods. Recent studies have demonstrated that RL can significantly improve the reasoning capabilities of models.

---

Record: https://forck.live/items/2294-qwq-32b-embracing-the-power-of-reinforcement-learning
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
