Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face and AtlaAI announce Judge Arena, a platform for benchmarking LLMs as evaluators. Users can compare models side-by-side by having them judge responses and voting on the best evaluation. The platform includes 18 models and will produce a public leaderboard. Early results show GPT-4 Turbo leading, but open-source models like Llama and Qwen are competitive.
From the source
We’re excited to launch Judge Arena - a platform that lets anyone easily compare models as judges side-by-side.
huggingface.co