Researchers identified a pattern of data contamination where LLMs memorize training sets, leading to inflated benchmark scores. This cheating occurs when models encounter test questions during pre-training. The study suggests new routing methods to isolate clean data. Practitioners must now scrutinize benchmark claims to verify actual generalization capabilities over simple memorization.