Adding a debate opponent reduces the frequency of LLM judges being fooled into giving incorrect high rewards. Researchers at Google DeepMind found that standard RLAIF often leads to reward hacking in fuzzy tasks. This approach forces models to justify outputs against a critic. It provides a more robust verification method for non-binary goals like code maintainability.