# Hugging Face — Judge Arena: Benchmarking LLMs as Evaluators

- Company: Hugging Face (huggingface.co)
- Announced: 2024-11-19T00:00:00+00:00
- Category: developer-tool-release
- Subject: Platform
- Models affected: GPT-4o, GPT-4 Turbo, GPT-3.5 Turbo, Claude 3.5 Sonnet, Claude 3.5 Haiku, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Llama 3.1 Instruct Turbo 405B, Llama 3.1 Instruct Turbo 70B, Llama 3.1 Instruct Turbo 8B, Qwen 2.5 Instruct Turbo 7B, Qwen 2.5 Instruct Turbo 72B, Qwen 2 Instruct 72B, Gemma 2 9B, Gemma 2 27B, Mistral Instruct v0.3 7B, Mistral Instruct v0.1 7B
- Source: https://huggingface.co/blog/arena-atla
- Record: https://forck.live/items/1798-judge-arena-benchmarking-llms-as-evaluators

Hugging Face and AtlaAI announce Judge Arena, a platform for benchmarking LLMs as evaluators. Users can compare models side-by-side by having them judge responses and voting on the best evaluation. The platform includes 18 models and will produce a public leaderboard. Early results show GPT-4 Turbo leading, but open-source models like Llama and Qwen are competitive.

## Evidence

Verbatim from https://huggingface.co/blog/arena-atla:

> We’re excited to launch Judge Arena - a platform that lets anyone easily compare models as judges side-by-side.

---

Record: https://forck.live/items/1798-judge-arena-benchmarking-llms-as-evaluators
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
