Researchers trained a new token on data generated via persona steering to test model self-interpretation. While the model expressed the intended traits, its explanations often diverged, labeling an "evil" vector as "dread." This gap suggests LLMs struggle to accurately describe their own internal steering mechanisms. Practitioners should distrust model-generated explanations of latent behaviors.