# Together AI — Together Evaluations now supports comparing top commercial APIs vs. open source models

- Company: Together AI (together.ai)
- Announced: 2026-02-02T00:00:00+00:00
- Category: capability-change
- Subject: Inference platform
- Models affected: gpt-5, gpt-5.2, claude-sonnet-4-5, claude-haiku-4-5, claude-opus-4-5, gemini-2.5-pro, gemini-2.5-flash, GPT-OSS 120B, Qwen3 235B
- Source: https://www.together.ai/blog/together-evaluations-v2
- Record: https://forck.live/items/2385-together-evaluations-now-supports-comparing-top-commercial-apis-vs-open-source

Together Evaluations now supports proprietary models from OpenAI, Anthropic, and Google as both judges and targets, allows evaluation of Together fine-tuned models, and provides new recipes demonstrating fine-tuned open judges that outperform GPT-5.2 at lower cost and higher speed, plus prompt optimization with GEPA.

## Evidence

Verbatim from https://www.together.ai/blog/together-evaluations-v2:

> Today, we’re excited to announce support for closed-source frontier models—including those from OpenAI, Anthropic, and Google—as both the judge and target model for rigorous cross-model benchmarking.

---

Record: https://forck.live/items/2385-together-evaluations-now-supports-comparing-top-commercial-apis-vs-open-source
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
