Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
ServiceNow and Hugging Face open-source PipelineRL, an experimental reinforcement learning implementation that addresses the trade-off between inference throughput and on-policy data collection by using inflight weight updates during RL training. It enables high inference throughput while keeping data on-policy, achieving competitive results on reasoning benchmarks compared to Open-Reasoner-Zero when training 7B and 32B models.
From the source
We are excited to open-source PipelineRL, an experimental RL implementation that tackles a fundamental challenge in large-scale Reinforcement Learning with LLMs: the trade-off between inference throughput and on-policy data collection.
huggingface.co