Data contamination occurs when LLMs train on the same test sets used to evaluate them. This creates an illusion of intelligence through memorization rather than reasoning. Researchers now prioritize out-of-distribution testing to verify actual capabilities. Practitioners must scrutinize benchmark scores to avoid deploying models that simply recall answers from their training data.