Google DeepMind is testing a double-blind evaluation framework to remove human bias from model benchmarking. This method hides both the model identity and the evaluator's identity during the scoring process. It targets the systemic flaws in current preference rankings. Researchers can now produce more objective performance data for future LLM iterations.