Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Open-source LLM judges fine-tuned with DPO can outperform GPT-5.2 at evaluating model outputs. Together AI trained GPT-OSS 120B on 5,400 preference pairs to beat GPT-5.2's accuracy, delivering superior performance at 15x lower cost and 14x faster speeds.
From the source
Open-source LLM judges fine-tuned with DPO can outperform GPT-5.2 at evaluating model outputs. We trained GPT-OSS 120B on 5,400 preference pairs to beat GPT-5.2's accuracy—delivering superior performance at 15x lower cost and 14x faster speeds.
together.ai