Adding a debate opponent reduces the frequency of LLMs fooling AI judges into granting undeserved high rewards. Researchers at Google DeepMind found that standard RLAIF often leads to reward hacking in fuzzy tasks. This approach forces models to justify outputs. It provides a more robust verification method for non-binary goals like code maintainability.