# Google Research — Building better AI benchmarks: How many raters are enough?

- Company: Google Research (research.google)
- Announced: 2026-03-31T16:16:00+00:00
- Category: research-paper
- Subject: Research
- Source: https://research.google/blog/building-better-ai-benchmarks-how-many-raters-are-enough/
- Record: https://forck.live/items/1351-building-better-ai-benchmarks-how-many-raters-are-enough

Google Research introduces an evaluation framework for ML models that optimizes the trade-off between the number of items and raters per item to build reproducible AI benchmarks that capture human disagreement. The research finds that the common practice of 3-5 raters is often insufficient and that more than 10 raters per item is often needed. The simulator has been open-sourced on GitHub.

## Evidence

Verbatim from https://research.google/blog/building-better-ai-benchmarks-how-many-raters-are-enough/:

> We introduce an evaluation framework for ML models, based on “gold” ratings data, that optimizes the trade-off between the number of items and raters per item, providing a roadmap for building highly reproducible AI benchmarks that capture the nuance of human disagreement.

---

Record: https://forck.live/items/1351-building-better-ai-benchmarks-how-many-raters-are-enough
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
