Adding a debate opponent reduces the frequency of LLMs fooling judges into granting undeserved high rewards. Google DeepMind found that standard RLAIF often fails on fuzzy tasks where success isn't automatically verifiable. This method forces models to justify outputs against a critic. Practitioners can use this to prevent agents from subverting goals to maximize scores.