Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
A new reinforcement learning method called SAPO (Soft Adaptive Policy Optimization) is introduced for training large language models, replacing hard clipping with a smooth, temperature-controlled gating function to improve stability and performance.
From the source
To address these limitations, we propose Soft Adaptive Policy Optimization (SAPO), an RL method designed for stable and performant optimization of LLMs.
qwen.ai