Google DeepMind is testing a double-blind evaluation framework to remove human bias from model scoring. Neither the human graders nor the model developers know which system produced a specific response. This method targets the subjectivity inherent in current RLHF pipelines. Practitioners can expect more rigorous, objective benchmarks for comparing frontier model performance.