# Together AI — Fine-tuning open LLM judges to outperform GPT-5.2

- Company: Together AI (together.ai)
- Announced: 2026-02-02T00:00:00+00:00
- Category: research-paper
- Subject: Inference platform
- Models affected: GPT-5.2, gpt-oss 120B, Qwen3 235B, Llama 4 Mav
- Pricing: $0.15 input / $0.60 output
- Source: https://www.together.ai/blog/fine-tuning-open-llm-judges-to-outperform-gpt-5-2
- Record: https://forck.live/items/2384-fine-tuning-open-llm-judges-to-outperform-gpt-5-2

Open-source LLM judges fine-tuned with DPO can outperform GPT-5.2 at evaluating model outputs. Together AI trained GPT-OSS 120B on 5,400 preference pairs to beat GPT-5.2's accuracy, delivering superior performance at 15x lower cost and 14x faster speeds.

## Evidence

Verbatim from https://www.together.ai/blog/fine-tuning-open-llm-judges-to-outperform-gpt-5-2:

> Open-source LLM judges fine-tuned with DPO can outperform GPT-5.2 at evaluating model outputs. We trained GPT-OSS 120B on 5,400 preference pairs to beat GPT-5.2's accuracy—delivering superior performance at 15x lower cost and 14x faster speeds.

---

Record: https://forck.live/items/2384-fine-tuning-open-llm-judges-to-outperform-gpt-5-2
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
