Researchers found that Llama-3.2-3B function vectors can pass behavioral and causal checks while performing the wrong task. These vectors appeared to shift months by k, but actually outputted adjacent months regardless of k. This failure occurs when few-shot prompts lack diversity. Practitioners cannot rely on standard validation metrics to guarantee model interpretability.