A new audit framework identifies how rater stress and distress create structured confounders in RLHF preference data. These state-dependent shifts differ from random noise and propagate directly through reward modeling. Researchers at arXiv prove that annotator mood encodes bias into policy optimization. Practitioners must now account for rater state to ensure objective model alignment.