Google DeepMind is testing double-blind evaluations to remove human bias from AI benchmarking. This method hides both the model identity and the evaluator's identity during testing. It prevents brand prestige from inflating scores. Researchers now have a cleaner metric to determine if a model actually performs better or just looks better.