Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog post provides a tutorial on reproducing DeepSeek R1's 'aha moment' using Group Relative Policy Optimization (GRPO) and the Countdown Game, training an open model with reinforcement learning.
From the source
In this blog post we want to recreate the small "aha moment" of DeepSeek-R1 using Group Relative Policy Optimization (GRPO) and the Countdown Game.
huggingface.co