Recent debates center on whether LLMs actually reason or simply mimic patterns from training data. Researchers argue that correct answers often stem from statistical shortcuts rather than logical deduction. This distinction matters for reliability. Practitioners must verify if a model's logic holds under slight perturbations or if it merely predicts the most likely token.