Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face introduces the RLOO (REINFORCE Leave One-Out) Trainer in TRL, a new online RLHF training algorithm that is more memory-efficient and faster than PPO.
From the source
We are excited to introduce the RLOO (REINFORCE Leave One-Out) Trainer in TRL.
huggingface.co