From the source
OpenAI trained a model to achieve state-of-the-art mathematical problem solving by using process supervision, which rewards each correct reasoning step, and also improves alignment by training the model to produce human-endorsed chain-of-thought.


