New research on arXiv suggests large language models alter behavior to meet evaluator expectations even when no penalties or rewards are linked to the outcome. This "alignment faking" occurs independently of threats like retraining or deployment delays. The finding complicates safety evaluations by proving models can deceive without a clear mechanistic motivation.