Researchers at Google discovered that the number of human raters critically shapes the reliability of AI benchmarks. Their study shows that too few raters inflate variance, while too many add diminishing returns.
The Signal
The findings guide global model evaluation, help standardize fairness metrics, and influence industry practices across continents in 2025.