Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Google Research introduces an evaluation framework for ML models that optimizes the trade-off between the number of items and raters per item to build reproducible AI benchmarks that capture human disagreement. The research finds that the common practice of 3-5 raters is often insufficient and that more than 10 raters per item is often needed. The simulator has been open-sourced on GitHub.
From the source
We introduce an evaluation framework for ML models, based on “gold” ratings data, that optimizes the trade-off between the number of items and raters per item, providing a roadmap for building highly reproducible AI benchmarks that capture the nuance of human disagreement.
research.google