Open-R1 project provides a one-week update on reproducing DeepSeek-R1's training pipeline and synthetic data, including evaluation results on MATH-500, integration of GRPO into TRL, and efforts to…
This blog post provides a tutorial on reproducing DeepSeek R1's 'aha moment' using Group Relative Policy Optimization (GRPO) and the Countdown Game, training an open model with reinforcement learning.
A newsletter looking back at key milestones, tools, and breakthroughs in AI and arts from 2024, and forward to 2025, highlighting open source releases in image and video generation.
A guide on deploying and fine-tuning DeepSeek R1 models on AWS using Hugging Face Inference Endpoints, Amazon Bedrock Marketplace, and Amazon SageMaker AI, including code snippets and hardware…
Hugging Face launches the Open-R1 project to systematically reconstruct DeepSeek-R1's data and training pipeline, aiming to provide transparency on how reinforcement learning can enhance reasoning…
Hugging Face announces the integration of serverless inference providers (fal, Replicate, Sambanova, Together AI) directly on Hub model pages and client SDKs, allowing users to run inference on a…
Hugging Face publishes a blog post reviewing open video generation models and outlining Diffusers team plans for supporting them.