Melanie Mitchell argues that current benchmarks fail to distinguish between genuine reasoning and sophisticated pattern matching. LLMs often mimic the structure of logic without understanding the underlying concepts. This gap in measurement makes it difficult for practitioners to determine when a model is actually solving a problem versus hallucinating a plausible answer.