Melanie Mitchell of the Santa Fe Institute argues that current evaluation methods fail to distinguish between genuine reasoning and sophisticated pattern matching. LLMs often produce text that mimics logic without understanding the underlying concepts. This gap in measurement makes it difficult to determine when these models can be trusted without human supervision.