Researchers propose a symbolic‑mechanistic evaluation framework that exposes when language models rely on memorization rather than generalization. By combining task‑specific rules with mechanistic interpretability, the method delivers pass/fail scores that pinpoint model weaknesses. Applied to NL‑to‑SQL, the approach uncovered claims, prompting a reassessment of evaluation standards.
The Signal
The study, published on ArXiv, signals a shift toward AI benchmarks worldwide.