# Hugging Face — Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial

- Company: Hugging Face (huggingface.co)
- Announced: 2025-01-31T10:29:40+00:00
- Subject: Platform
- Models affected: DeepSeek-R1, DeepSeek-R1-Zero, Qwen2.5-3B-Instruct
- Source: https://huggingface.co/blog/open-r1/mini-r1-contdown-game
- Record: https://forck.live/items/1757-mini-r1-reproduce-deepseek-r1-aha-moment-a-rl-tutorial

This blog post provides a tutorial on reproducing DeepSeek R1's 'aha moment' using Group Relative Policy Optimization (GRPO) and the Countdown Game, training an open model with reinforcement learning.

## Evidence

Verbatim from https://huggingface.co/blog/open-r1/mini-r1-contdown-game:

> In this blog post we want to recreate the small "aha moment" of DeepSeek-R1 using Group Relative Policy Optimization (GRPO) and the Countdown Game.

---

Record: https://forck.live/items/1757-mini-r1-reproduce-deepseek-r1-aha-moment-a-rl-tutorial
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
