Researchers identified a pattern of data leakage where models memorize training sets. This contamination skews benchmarks, making LLMs appear smarter than they are. The study highlights a systemic failure in how developers curate evaluation data. Practitioners must now implement stricter data scrubbing to ensure benchmark results reflect actual reasoning rather than rote memorization.