Paper arXiv:2608.12325v1 claims that current generative AI lacks a concrete operational definition for reasoning. This ambiguity makes evaluating construct validity impossible and hinders progress toward trustworthy systems. The authors propose returning to symbolic AI logic to create verifiable benchmarks. Practitioners should expect a push for more rigorous, rule-based evaluation metrics.