Researchers identified a pattern of data leakage where LLMs memorize training sets, leading to inflated benchmark scores. This contamination occurs when test questions appear in the training data. Developers must now implement stricter data filtering. This ensures that model performance reflects actual reasoning rather than simple rote memorization of existing answers.