Researchers propose a mechanism‑aware evaluation that blends symbolic rules with mechanistic interpretability, revealing where models truly generalize versus exploit shortcuts. Tested on NL‑to‑SQL tasks, the method shows that models lacking schema knowledge can achieve high accuracy by memorization, while grounded models reveal genuine understanding.
The Signal
This shift could reshape global AI benchmarking and foster more reliable deployments.