Stressful annotation conditions create structured confounds in pairwise preference labels. This rater state shift differs from random noise, as it encodes the annotator's emotional state rather than output quality. These biases propagate through reward models into final policy optimization. Researchers now have a framework to audit these hidden biases in RLHF pipelines.